用视频生成模型做视觉编码,让单图模型也能推理多帧动态。
Can World Models Benefit VLMs for World Dynamics?
- 把视频扩散模型改造成生成编码器,单步去噪得视觉嵌入。
- 新模型在空间推理上超越基线,单图可完成多帧理解。
- 适合追求动态感知的通用视觉语言模型研究者。
基于互联网规模视频数据训练的生成式世界模型,正被视为强大的世界模拟器,能生成结构、运动和物理上一致且合理的动态。这引发了一个问题:随着强视频基础模型的出现,它们是否能取代传统视觉编码范式,实现通用多模态理解?尽管近期研究开始探索世界模型在常见视觉任务中的潜力,但这些探索通常缺乏对通用多模态任务的系统性考察。本文旨在研究将世界模型先验迁移到视觉语言模型(VLM)的能力:我们重新利用一个视频扩散模型作为生成编码器,执行单步去噪,并将得到的潜在表示视为一组视觉嵌入。我们实证研究了这类模型,称之为世界-语言模型(WorldLMs),发现生成编码器能捕捉对下游理解有用的潜在表示,与传统编码器有明显区别。我们提出的最优变体称为动态视觉对齐器(DyVA),进一步发现该方法显著提升了空间推理能力,使单图模型具备多帧推理能力。通过精心设计的一系列视觉推理任务,我们发现DyVA超越了开源和专有基线,达到或接近当前最佳性能。我们归因于世界模型从视频预训练中继承的运动一致性内化。最后,我们系统性地探索了多种模型设计,指出了未来工作的有希望方向。我们希望本研究能为一类利用世界模型先验的新一代VLM铺平道路,并迈向通用视觉学习者的前景。
原文摘要 · Abstract (English)
Trained on internet-scale video data, generative world models are increasingly recognized as powerful world simulators that can generate consistent and plausible dynamics over structure, motion, and physics. This raises a natural question: with the advent of strong video foundational models, might they supplant conventional vision encoder paradigms for general-purpose multimodal understanding? While recent studies have begun to explore the potential of world models on common vision tasks, these explorations typically lack a systematic investigation of generic, multimodal tasks. In this work, we strive to investigate the capabilities when world model priors are transferred into Vision-Language Models: we re-purpose a video diffusion model as a generative encoder to perform a single denoising step and treat the resulting latents as a set of visual embedding. We empirically investigate this class of models, which we refer to as World-Language Models (WorldLMs), and we find that generative encoders can capture latents useful for downstream understanding that show distinctions from conventional encoders. Naming our best-performing variant Dynamic Vision Aligner (DyVA), we further discover that this method significantly enhances spatial reasoning abilities and enables single-image models to perform multi-frame reasoning. Through the curation of a suite of visual reasoning tasks, we find DyVA to surpass both open-source and proprietary baselines, achieving state-of-the-art or comparable performance. We attribute these gains to WorldLM's inherited motion-consistency internalization from video pre-training. Finally, we systematically explore extensive model designs to highlight promising directions for future work. We hope our study can pave the way for a new family of VLMs that leverage priors from world models and are on a promising path towards generalist vision learners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。