Skip to main content
Cornell University
Learn about arXiv becoming an independent nonprofit.
We gratefully acknowledge support from the Simons Foundation, member institutions, and all contributors. Donate
arxiv logo > cs > arXiv:2509.20328

Help | Advanced Search

Computer Science > Machine Learning

(cs)
[Submitted on 24 Sep 2025 (v1), last revised 29 Sep 2025 (this version, v2)]

Title:Video models are zero-shot learners and reasoners

Authors:Thaddäus Wiedemer, Yuxuan Li, Paul Vicol, Shixiang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini, Robert Geirhos
View a PDF of the paper titled Video models are zero-shot learners and reasoners, by Thadd\"aus Wiedemer and 8 other authors
View PDF HTML (experimental)
Abstract:The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models. This transformation emerged from simple primitives: large, generative models trained on web-scale data. Curiously, the same primitives apply to today's generative video models. Could video models be on a trajectory towards general-purpose vision understanding, much like LLMs developed general-purpose language understanding? We demonstrate that Veo 3 can solve a broad variety of tasks it wasn't explicitly trained for: segmenting objects, detecting edges, editing images, understanding physical properties, recognizing object affordances, simulating tool use, and more. These abilities to perceive, model, and manipulate the visual world enable early forms of visual reasoning like maze and symmetry solving. Veo's emergent zero-shot capabilities indicate that video models are on a path to becoming unified, generalist vision foundation models.
Comments: Project page: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
Cite as: arXiv:2509.20328 [cs.LG]
  (or arXiv:2509.20328v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2509.20328
arXiv-issued DOI via DataCite

Submission history

From: Robert Geirhos [view email]
[v1] Wed, 24 Sep 2025 17:17:27 UTC (29,344 KB)
[v2] Mon, 29 Sep 2025 20:44:46 UTC (29,332 KB)
Full-text links:

Access Paper:

    View a PDF of the paper titled Video models are zero-shot learners and reasoners, by Thadd\"aus Wiedemer and 8 other authors
  • View PDF
  • HTML (experimental)
  • TeX Source
view license

Current browse context:

cs.LG
< prev   |   next >
new | recent | 2025-09
Change to browse by:
cs
cs.AI
cs.CV
cs.RO

References & Citations

  • NASA ADS
  • Google Scholar
  • Semantic Scholar
Loading...

Bookmark

BibSonomy Reddit

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)
  • About
  • Help
  • contact arXivClick here to contact arXiv Contact
  • subscribe to arXiv mailingsClick here to subscribe Subscribe
  • Copyright
  • Privacy Policy
  • Web Accessibility Assistance
  • arXiv Operational Status