arXiv:2512.14691cs.CLcs.CV2025-12被引 9

评测视频生成模型的推理能力,发现它们在逻辑和空间规划上严重不足。

MMGR: Multi-Modal Generative Reasoning

  • 构建五维推理能力评估框架,覆盖物理、逻辑、时空等维度。
  • 模型在抽象推理任务中准确率低于10%,长时空间规划能力差。
  • 适合关注生成模型真实世界推理能力的研究者和开发者。

视频基础模型能生成视觉逼真且时间连贯的内容,但其作为世界模拟器的可靠性取决于是否捕捉了物理、逻辑和空间约束。现有指标如弗雷切特视频距离(FVD)侧重感知质量,忽视因果性、物理规律和全局一致性等推理错误。我们提出MMGR(多模态生成推理评估与基准),基于五大推理能力:物理、逻辑、三维空间、二维空间和时间。评估涵盖三个领域:抽象推理(ARC-AGI、数独)、具身导航(真实3D环境下的导航与定位)、物理常识(运动及组合交互)。MMGR采用细粒度指标,要求视频与图像生成整体正确。我们对领先视频模型(Veo-3、Sora-2、Wan-2.2)和图像模型(Nano-banana、Nano-banana Pro、GPT-4o-image、Qwen-image)进行测评,发现各领域表现差异显著。模型在物理常识任务中表现中等,但在抽象推理任务(如ARC-AGI)中准确率低于10%,在具身场景中的长时空间规划能力严重不足。分析表明当前模型存在过度依赖感知数据、全局状态不一致、目标偏重视觉合理性而非因果正确性等关键缺陷。MMGR提供统一诊断基准,推动具备推理能力的生成世界模型发展。

原文摘要 · Abstract (English)

Video foundation models generate visually realistic and temporally coherent content, but their reliability as world simulators depends on whether they capture physical, logical, and spatial constraints. Existing metrics such as Frechet Video Distance (FVD) emphasize perceptual quality and overlook reasoning failures, including violations of causality, physics, and global consistency. We introduce MMGR (Multi-Modal Generative Reasoning Evaluation and Benchmark), a principled evaluation framework based on five reasoning abilities: Physical, Logical, 3D Spatial, 2D Spatial, and Temporal. MMGR evaluates generative reasoning across three domains: Abstract Reasoning (ARC-AGI, Sudoku), Embodied Navigation (real-world 3D navigation and localization), and Physical Commonsense (sports and compositional interactions). MMGR applies fine-grained metrics that require holistic correctness across both video and image generation. We benchmark leading video models (Veo-3, Sora-2, Wan-2.2) and image models (Nano-banana, Nano-banana Pro, GPT-4o-image, Qwen-image), revealing strong performance gaps across domains. Models show moderate success on Physical Commonsense tasks but perform poorly on Abstract Reasoning (below 10 percent accuracy on ARC-AGI) and struggle with long-horizon spatial planning in embodied settings. Our analysis highlights key limitations in current models, including overreliance on perceptual data, weak global state consistency, and objectives that reward visual plausibility over causal correctness. MMGR offers a unified diagnostic benchmark and a path toward reasoning-aware generative world models.

视频生成推理评估多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。