Stage 3 · Expert · D11
Seeing and Hearing Together
To a transformer, a photo, a voice note and a sentence are the same thing: a line of tokens.
14 lessons · 135 minDemanding
About this chapter
A model that reads your menu photo or answers your voice note never sees or hears anything. It reads lines of tokens. This chapter shows how a picture, a document, a sound and a video each become such a line, how pictures and words learn to meet in one space, and how the pieces are wired into one language model. It ends where these models fail, how to measure them honestly, and what happens when the output is a robot arm moving.
What you will be able to do
- 19 min
Every Sense Becomes a Line
Say how a picture becomes a line of tokens, and count what it costs.
- 28 min
Pictures Meet Their Captions
Explain how photos and captions end up in one shared space without anyone labeling them.
- 310 min
Right Pairs Up, Wrong Pairs Down
Read a batch's grid of scores, and say what the contrastive loss asks of each row.
- 49 min
Classify by Writing Captions
Sort photos into classes nobody trained on, and say where it slips.
- 510 min
Pictures Inside a Language Model
Trace a photo through an encoder and a projector into a language model, and say which parts learn when.
- 69 min
What a Picture Costs
Estimate the tokens a picture costs, and name two ways models keep that bill down.
- 711 min
Reading a Page
Compare a step-by-step reader with one that reads straight from pixels, and say what makes Arabic script hard.
- 810 min
Sound as a Picture
Read a spectrogram, say how it is made from sound, and why speech models squeeze its high pitches.
- 910 min
From Voice to Words
Trace speech through a recognizer, score a transcript by word error rate, and say why Arabic is harder to score.
- 1010 min
Words Back to Voice
Trace text to speech through a spectrogram or through sound tokens, and count what a second of speech costs.
- 1110 min
Pictures in Time
Price a video in tokens, choose how often to sample it, and say what sampling can miss.
- 128 min
Making, Not Only Reading
Name the two ways a model makes a picture or a sound, and count what a picture costs as tokens.
- 1310 min
Where They Fail
Spot a model that invents objects, grade it on a fair test, and name the note that fools a picture reader.
- 1411 min
Models That Act
Write a robot's move as tokens and back, and say what vision-language-action models borrow from the web and what they still lack.