arXiv:2511.15065cs.CVcs.AI2025-11被引 16

用迷宫任务测试视频模型的推理能力,发现其空间推理表现优于视觉语言模型。

Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks

  • 基于迷宫求解设计视频推理评测基准,涵盖7920个生成视频。
  • 微调后视频模型在空间推理上超越主流视觉语言模型,复杂任务泛化能力强。
  • 推理时多样采样可提升10%-20%可靠性,适合空间规划类应用研究者。

视频模型在高质量视频生成与连贯运动动态方面取得显著进展。类比语言模型从文本生成到文本推理的发展,视频模型的发展促使我们思考:视频模型能否通过视频生成进行推理?相比离散的文本语料,视频具备明确的空间布局与时间连续性,是理想的空间推理载体。本文探索视频推理范式,提出VR-Bench——一个系统评估视频模型推理能力的综合性基准。该基准基于需空间规划与多步推理的迷宫求解任务,包含5种迷宫类型和多样视觉风格的7,920个程序生成视频。实证分析表明,SFT能有效激发视频模型的推理能力;视频模型在推理中展现出更强的空间感知,优于领先视觉语言模型,并在不同场景、任务及复杂度下具有良好泛化性。进一步发现,推理时采用多样化采样可使推理可靠性提升10%-20%。这些结果凸显了视频推理在空间推理任务中的独特潜力与可扩展性。

原文摘要 · Abstract (English)

Video Models have achieved remarkable success in high-fidelity video generation with coherent motion dynamics. Analogous to the development from text generation to text-based reasoning in language modeling, the development of video models motivates us to ask: Can video models reason via video generation? Compared with the discrete text corpus, video grounds reasoning in explicit spatial layouts and temporal continuity, which serves as an ideal substrate for spatial reasoning. In this work, we explore the reasoning via video paradigm and introduce VR-Bench -- a comprehensive benchmark designed to systematically evaluate video models' reasoning capabilities. Grounded in maze-solving tasks that inherently require spatial planning and multi-step reasoning, VR-Bench contains 7,920 procedurally generated videos across five maze types and diverse visual styles. Our empirical analysis demonstrates that SFT can efficiently elicit the reasoning ability of video model. Video models exhibit stronger spatial perception during reasoning, outperforming leading VLMs and generalizing well across diverse scenarios, tasks, and levels of complexity. We further discover a test-time scaling effect, where diverse sampling during inference improves reasoning reliability by 10--20%. These findings highlight the unique potential and scalability of reasoning via video for spatial reasoning tasks.

视频推理空间规划迷宫求解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。