D5 · Seeing
When Attention Came for Vision
Compare a vision transformer with a convolutional network on built-in assumptions and hunger for data.
In 2020 a Google team tried something odd. They cut a photo into 16 by 16 pixel squares and treated each square like a word. The line of squares went into a transformer, the model built for sentences. No convolution at all. The paper's title: An Image Is Worth 16x16 Words.