D11 · Seeing and Hearing Together
Pictures Inside a Language Model
Trace a photo through an encoder and a projector into a language model, and say which parts learn when.
CLIP can match a photo to a sentence, but it cannot answer a question about it. It has no voice. A language model has a voice and no eyes. So plug one into the other.