Stage 3 · Expert · A1
Building With Model APIs
Most AI products are two hundred lines of glue around one API call.
11 lessons · 107 minSteady
About this chapter
The call itself is the easy part. What separates a demo from a product is a key nobody can copy, streaming that feels instant, retries that do not double your bill, and an eval set that blocks a merge before a prompt change reaches anyone. This chapter builds all of it around one small support feature, prices every request in cents, and turns a wait into a number you can plan with. By the end you choose a model with your own test instead of a leaderboard.
What you will be able to do
- 19 min
Anatomy of a Request
Read a request and response, field by field, including token usage.
- 28 min
Where the Key Lives
Keep the provider key off every device and put a spend cap in front of it.
- 310 min
Streaming
Stream a response to the interface and handle partial and interrupted output.
- 49 min
Structured Outputs
Validate model output against a schema and recover when it does not match.
- 511 min
Tool Calling in Production
Wire a tool end to end, including the result round trip and errors.
- 610 min
Prompt Caching
Order a prompt so the stable prefix is cacheable and measure the saving.
- 710 min
Cost and Latency Budgets
Set a per-request budget and measure where the time actually goes.
- 810 min
Retries and Fallbacks
Retry safely with backoff and fall back without breaking the user's flow.
- 910 min
Seeing What Happened
Log prompts, versions, latency and cost in a way you can query later.
- 1011 min
Evals in CI
Write a small eval suite that blocks a merge when quality drops.
- 119 min
Choosing a Model
Pick a model with your own eval, your latency target and your budget.