Description
The authors explore how a modern generative video model, Veo 3, demonstrates a surprising level of zero-shot capability across many visual tasks it never explicitly trained for. Tasks include segmentation, edge detection, image editing, understanding physical dynamics, affordance recognition, and even reasoning puzzles like maze solving and symmetry completion. They argue that these emergent abilities suggest video models are evolving toward generalist vision foundation models, much like how LLMs transformed language understanding. Their experiments support a hierarchy of visual competence — from perception to modeling to manipulation to reasoning — showing that Veo 3 can “see, model, act, and reason” with visual content in ways that hint at future unified vision systems.