视频预训练让视觉模型更高效,比语言模型更适合解决视觉问题。
Rethinking Visual Intelligence: Insights from Video Pretraining
- 用视频扩散模型在时空数据上预训练,引入结构与动态的先验知识。
- 在多个视觉任务中,视频模型仅用少量样本就达到更高准确率。
- 适合追求视觉通用智能、注重样本效率的研究者或开发者。
大型语言模型(LLMs)在大规模预训练后能以极少监督快速适应新任务,但这一成功尚未有效延伸至视觉领域,现有模型在组合理解、样本效率和通用问题求解方面仍表现不佳。本文探讨视频扩散模型(VDMs)作为弥合差距的潜力:在时空数据上预训练使模型具备强结构与动态的归纳偏置,我们假设这可支持广泛的任务适应能力。为此,设计了控制实验,让预训练的LLM和预训练的VDM分别搭配轻量适配器,在各自自然模态下处理任务。在包括ARC-AGI、ConceptARC、视觉游戏、路径规划和细胞自动机等基准测试中,VDMs展现出显著更高的数据效率。结果表明,视频预训练提供的归纳偏置有助于推动视觉基础模型的发展。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated that large-scale pretraining enables systems to adapt rapidly to new problems with little supervision in the language domain. This success, however, has not translated as effectively to the visual domain, where models, including LLMs, continue to struggle with compositional understanding, sample efficiency, and general-purpose problem-solving. We investigate Video Diffusion Models (VDMs) as a promising direction for bridging this gap. Pretraining on spatiotemporal data endows these models with strong inductive biases for structure and dynamics, which we hypothesize can support broad task adaptability. To test this, we design a controlled evaluation in which both a pretrained LLM and a pretrained VDM are equipped with lightweight adapters and presented with tasks in their natural modalities. Across benchmarks including ARC-AGI, ConceptARC, visual games, route planning, and cellular automata, VDMs demonstrate higher data efficiency than their language counterparts. Taken together, our results indicate that video pretraining offers inductive biases that support progress toward visual foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。