arXiv:2511.13853cs.CV2025-11被引 14

首个评估视频模型推理能力的基准,揭示视觉生成不等于真实思维。

Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark

  • 将视觉推理分解为6大认知维度、24个子任务,构建可量化的评估框架。
  • 发现顶尖视频模型在视觉质量出色但推理深度严重不足,存在巨大差距。
  • 适合研究世界模拟器、具身智能与多步推理的开发者和学者使用。

尽管链式思维(CoT)提示使大语言模型具备复杂符号推理能力,但其仍局限于离散文本,无法模拟现实世界连续且受物理规律支配的动态过程。近期视频生成模型通过链式帧(CoF)推理——将思维转化为逐帧视觉序列,每帧代表一个物理基础的推理步骤——展现出成为世界模拟器的潜力。然而,现有评估基准聚焦于生成质量或对齐度,未涵盖CoF推理能力,难以衡量模型在多步规划、算法逻辑或抽象模式外推等方面的认知水平。这一评估空白阻碍了对模型能力的系统理解与改进指导。为此,我们提出Gen-ViRe(生成式视觉推理基准),基于认知科学与真实应用场景,将CoF推理拆解为六项认知维度及24个子任务。通过多源数据构建、极简提示协议与混合视觉语言模型辅助评估,实现对视频模型作为推理者的能力首次定量评估。对前沿系统的实验表明,视觉表现优异的模型在实际推理深度上存在显著不足,确立了基线与诊断工具,推动真正世界模拟器的发展。

原文摘要 · Abstract (English)

While Chain-of-Thought (CoT) prompting enables sophisticated symbolic reasoning in LLMs, it remains confined to discrete text and cannot simulate the continuous, physics-governed dynamics of the real world. Recent video generation models have emerged as potential world simulators through Chain-of-Frames (CoF) reasoning -- materializing thought as frame-by-frame visual sequences, with each frame representing a physically-grounded reasoning step. Despite compelling demonstrations, a challenge persists: existing benchmarks, focusing on fidelity or alignment, do not assess CoF reasoning and thus cannot measure core cognitive abilities in multi-step planning, algorithmic logic, or abstract pattern extrapolation. This evaluation void prevents systematic understanding of model capabilities and principled guidance for improvement. We introduce Gen-ViRe (Generative Visual Reasoning Benchmark), a framework grounded in cognitive science and real-world AI applications, which decomposes CoF reasoning into six cognitive dimensions -- from perceptual logic to abstract planning -- and 24 subtasks. Through multi-source data curation, minimal prompting protocols, and hybrid VLM-assisted evaluation with detailed criteria, Gen-ViRe delivers the first quantitative assessment of video models as reasoners. Our experiments on SOTA systems reveal substantial discrepancies between impressive visual quality and actual reasoning depth, establishing baselines and diagnostic tools to advance genuine world simulators.

视觉推理世界模拟器视频生成认知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。