Skip to content

Stage 3 · Expert · D11

Seeing and Hearing Together

To a transformer, a photo, a voice note and a sentence are the same thing: a line of tokens.

14 lessons · 135 minDemanding

About this chapter

A model that reads your menu photo or answers your voice note never sees or hears anything. It reads lines of tokens. This chapter shows how a picture, a document, a sound and a video each become such a line, how pictures and words learn to meet in one space, and how the pieces are wired into one language model. It ends where these models fail, how to measure them honestly, and what happens when the output is a robot arm moving.

What you will be able to do

  1. 1

    Every Sense Becomes a Line

    Say how a picture becomes a line of tokens, and count what it costs.

    9 min
  2. 2

    Pictures Meet Their Captions

    Explain how photos and captions end up in one shared space without anyone labeling them.

    8 min
  3. 3

    Right Pairs Up, Wrong Pairs Down

    Read a batch's grid of scores, and say what the contrastive loss asks of each row.

    10 min
  4. 4

    Classify by Writing Captions

    Sort photos into classes nobody trained on, and say where it slips.

    9 min
  5. 5

    Pictures Inside a Language Model

    Trace a photo through an encoder and a projector into a language model, and say which parts learn when.

    10 min
  6. 6

    What a Picture Costs

    Estimate the tokens a picture costs, and name two ways models keep that bill down.

    9 min
  7. 7

    Reading a Page

    Compare a step-by-step reader with one that reads straight from pixels, and say what makes Arabic script hard.

    11 min
  8. 8

    Sound as a Picture

    Read a spectrogram, say how it is made from sound, and why speech models squeeze its high pitches.

    10 min
  9. 9

    From Voice to Words

    Trace speech through a recognizer, score a transcript by word error rate, and say why Arabic is harder to score.

    10 min
  10. 10

    Words Back to Voice

    Trace text to speech through a spectrogram or through sound tokens, and count what a second of speech costs.

    10 min
  11. 11

    Pictures in Time

    Price a video in tokens, choose how often to sample it, and say what sampling can miss.

    10 min
  12. 12

    Making, Not Only Reading

    Name the two ways a model makes a picture or a sound, and count what a picture costs as tokens.

    8 min
  13. 13

    Where They Fail

    Spot a model that invents objects, grade it on a fair test, and name the note that fools a picture reader.

    10 min
  14. 14

    Models That Act

    Write a robot's move as tokens and back, and say what vision-language-action models borrow from the web and what they still lack.

    11 min

Before you start

Keep going

All chapters