Skip to content

Stage 3 · Expert · A1

Building With Model APIs

Most AI products are two hundred lines of glue around one API call.

11 lessons · 107 minSteady

About this chapter

The call itself is the easy part. What separates a demo from a product is a key nobody can copy, streaming that feels instant, retries that do not double your bill, and an eval set that blocks a merge before a prompt change reaches anyone. This chapter builds all of it around one small support feature, prices every request in cents, and turns a wait into a number you can plan with. By the end you choose a model with your own test instead of a leaderboard.

What you will be able to do

  1. 1

    Anatomy of a Request

    Read a request and response, field by field, including token usage.

    9 min
  2. 2

    Where the Key Lives

    Keep the provider key off every device and put a spend cap in front of it.

    8 min
  3. 3

    Streaming

    Stream a response to the interface and handle partial and interrupted output.

    10 min
  4. 4

    Structured Outputs

    Validate model output against a schema and recover when it does not match.

    9 min
  5. 5

    Tool Calling in Production

    Wire a tool end to end, including the result round trip and errors.

    11 min
  6. 6

    Prompt Caching

    Order a prompt so the stable prefix is cacheable and measure the saving.

    10 min
  7. 7

    Cost and Latency Budgets

    Set a per-request budget and measure where the time actually goes.

    10 min
  8. 8

    Retries and Fallbacks

    Retry safely with backoff and fall back without breaking the user's flow.

    10 min
  9. 9

    Seeing What Happened

    Log prompts, versions, latency and cost in a way you can query later.

    10 min
  10. 10

    Evals in CI

    Write a small eval suite that blocks a merge when quality drops.

    11 min
  11. 11

    Choosing a Model

    Pick a model with your own eval, your latency target and your budget.

    9 min

Before you start

Keep going

All chapters