用生成视频帧模拟思考过程,提升视觉推理能力
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
- 以生成视频帧作为推理中间步骤,替代传统文本逻辑
- 零样本泛化能力强,未见场景下仍表现稳定
- 增加生成帧数可提升复杂路径推理效果,适合构建智能体
视觉语言模型在文本推理上表现优异,但在细粒度空间理解与连续动作规划方面表现不佳,难以模拟复杂视觉推理所需动态过程。本文提出通过视频生成模型实现视觉推理,认为生成的帧可作为初始状态与解决方案之间的中间推理步骤。我们在两个不同场景下评估:迷宫导航(低视觉变化、离散序列规划)和拼图游戏(高视觉变化、连续操作)。实验揭示三个关键发现:(1) 强鲁棒零样本泛化能力,在两任务中均在未见数据分布上无需微调即表现良好;(2) 视觉上下文利用有效,模型能显式使用代理图标与拼图形状等视觉线索,保持高视觉一致性并适应未见模式;(3) 视觉测试时缩放效应明显,在序列规划中增加生成视频长度(即视觉推理预算)可显著提升对空间与时间复杂路径的零样本泛化能力。结果表明,视频生成不仅是媒体工具,更是一种可扩展、可泛化的视觉推理范式。
原文摘要 · Abstract (English)
Vision-Language Models have excelled at textual reasoning, but they often struggle with fine-grained spatial understanding and continuous action planning, failing to simulate the dynamics required for complex visual reasoning. In this work, we formulate visual reasoning by means of video generation models, positing that generated frames can act as intermediate reasoning steps between initial states and solutions. We evaluate their capacity in two distinct regimes: Maze Navigation for sequential discrete planning with low visual change and Tangram Puzzle for continuous manipulation with high visual change. Our experiments reveal three critical insights: (1) Robust Zero-Shot Generalization: In both tasks, the model demonstrates strong performance on unseen data distributions without specific finetuning. (2) Visual Context: The model effectively uses visual context as explicit control, such as agent icons and tangram shapes, enabling it to maintain high visual consistency and adapt its planning capability robustly to unseen patterns. (3) Visual Test-Time Scaling: We observe a test-time scaling law in sequential planning; increasing the generated video length (visual inference budget) empowers better zero-shot generalization to spatially and temporally complex paths. These findings suggest that video generation is not merely a media tool, but a scalable, generalizable paradigm for visual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。