Video models are zero-shot learners and reasoners

Paper
2025-09-28

Description

The authors explore how a modern generative video model, Veo 3, demonstrates a surprising level of zero-shot capability across many visual tasks it never explicitly trained for. Tasks include segmentation, edge detection, image editing, understanding physical dynamics, affordance recognition, and even reasoning puzzles like maze solving and symmetry completion. They argue that these emergent abilities suggest video models are evolving toward generalist vision foundation models, much like how LLMs transformed language understanding. Their experiments support a hierarchy of visual competence — from perception to modeling to manipulation to reasoning — showing that Veo 3 can “see, model, act, and reason” with visual content in ways that hint at future unified vision systems.
PDF Preview

User Reviews

No reviews yet for this resource.