To a transformer, a photo, a voice note and a sentence are the same thing: a line of tokens.
D11 · Seeing and Hearing Together
A model that reads your menu photo or answers your voice note never sees or hears anything. It reads lines of tokens. This chapter shows how a picture, a document, a sound and a video each become such a line, how pictures and words learn to meet in one space, and how the pieces are wired into one language model. It ends where these models fail, how to measure them honestly, and what happens when the output is a robot arm moving.
- Every Sense Becomes a Line
- Pictures Meet Their Captions
- Right Pairs Up, Wrong Pairs Down
- Classify by Writing Captions
- Pictures Inside a Language Model
- What a Picture Costs
- Reading a Page
- Sound as a Picture
- From Voice to Words
- Words Back to Voice
- Pictures in Time
- Making, Not Only Reading
- Where They Fail
- Models That Act