D11 · Seeing and Hearing Together
Models That Act
Write a robot's move as tokens and back, and say what vision-language-action models borrow from the web and what they still lack.
Tell a robot arm: put the banana in the bowl. It sees the table through a camera and reads your words. Then it must answer with movement, not with a sentence.