arXiv:2509.20328cs.LGcs.AI2025-09被引 204

视频模型无需训练即可完成多种视觉任务,展现类语言模型的通用能力。

Video models are zero-shot learners and reasoners

  • 基于大规模生成式训练,视频模型可零样本执行未预训练任务。
  • 在物体分割、边缘检测、物理属性理解等10+任务上表现优异。
  • 适合研究通用视觉智能与跨模态推理的学者参考。

大型语言模型(LLMs)的零样本能力推动自然语言处理从专用模型转向统一的通用基础模型。这一转变源于简单的范式:在海量网络数据上训练的大规模生成模型。有趣的是,同样的范式也适用于当今的生成式视频模型。视频模型是否正走上类似路径,发展出像LLMs一样的通用语言理解能力?我们证明,Veo 3能够在未显式训练的任务中解决广泛问题:包括物体分割、边缘检测、图像编辑、物理属性理解、物体功能识别、工具使用模拟等。这些感知、建模与操控视觉世界的能力,使模型初步具备迷宫求解、对称性分析等视觉推理能力。Veo的涌现式零样本能力表明,视频模型正朝着统一的通用视觉基础模型迈进。

原文摘要 · Abstract (English)

The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models. This transformation emerged from simple primitives: large, generative models trained on web-scale data. Curiously, the same primitives apply to today's generative video models. Could video models be on a trajectory towards general-purpose vision understanding, much like LLMs developed general-purpose language understanding? We demonstrate that Veo 3 can solve a broad variety of tasks it wasn't explicitly trained for: segmenting objects, detecting edges, editing images, understanding physical properties, recognizing object affordances, simulating tool use, and more. These abilities to perceive, model, and manipulate the visual world enable early forms of visual reasoning like maze and symmetry solving. Veo's emergent zero-shot capabilities indicate that video models are on a path to becoming unified, generalist vision foundation models.

视频生成零样本视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。