arXiv:2607.09024cs.CVcs.AI2026-07被引 6

用视频生成预训练,让模型学会通用视觉理解。

Video Generation Models are General-Purpose Vision Learners

论文配图:Video Generation Models are General-Purpose Vision Learners
图 1 · 摘自论文原文
  • 以视频生成模型为骨干,通过文本指令完成多种视觉任务。
  • 在深度、法向、姿态等任务上超越或媲美专用模型,数据效率更高。
  • 合成视频训练的模型能泛化到真实世界和新物体,具涌现能力。

受下一个词预测驱动,自然语言处理从专用模型转向强大的通用基础模型。那么,计算机视觉如何实现类似突破?本文认为大规模文本到视频生成是计算机视觉的重要预训练范式,可提供时空先验、视觉-语言对齐及可扩展性,推动通用视觉智能发展。我们提出GenCeption,利用预训练的视频生成扩散模型构建前馈感知模型,可通过文本指令执行多种视觉任务。实验证明,GenCeption在深度估计、表面法向、相机位姿、表情指代分割、3D关键点预测等多样化任务中达到领先性能,常优于或媲美专用模型(如DepthAnything3、SAM3、D4RT、VGGT-Omega、Sapiens、David、Genmo、Lotus-2)。此外,在相同设置下,其视频生成预训练骨干优于V-JEPA和Video MAE等替代方案。重要的是,GenCeption展现出初步的数据与模型缩放特性,仅需7至500倍更少的训练数据,即可达到D4RT和VGGT-Omega等领先模型的性能。最后,该模型在仅用合成人类视频训练后,仍能泛化至真实场景和分布外类别(如动物、机器人),表现出有趣的涌现行为。这些发现表明,视频生成不仅是内容合成工具,更是通往物理世界通用视觉智能的基础路径。

原文摘要 · Abstract (English)

Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Omega, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency, where it achieves comparable performance with leading models like D4RT and VGGT-Omega with 7 to 500 less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world. Project page: https://genception.github.io

视频生成通用模型视觉理解扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。