R2 · Landmark Papers, Rebuilt
CLIP (2021)
Build the image and text matching matrix, and say what zero-shot really means.
A vision model was trained on a fixed list of classes. ImageNet had a thousand of them. Ask it about anything outside that list and there was nothing to do but collect labeled photos and train again.