arXiv:2602.08236cs.CVcs.AI2026-02被引 9

让AI智能判断何时想象、想象多少,提升视觉空间推理效率与准确率。

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

  • 基于世界模型动态判断是否需要视觉想象,按需调用
  • 在多个基准上实现比固定想象策略更高的准确率,减少60%以上模型调用
  • 无需人工标注,通过奖励机制自动学习想象策略,适合实际部署

尽管多模态大模型进展迅速,当正确答案依赖于未见视角下的场景外观时,视觉空间推理仍不可靠。近期工作通过引入世界模型进行视觉想象来增强推理,但何时需要想象、想象多少合适、何时会适得其反等问题仍不明确。实践中盲目想象会增加计算开销,甚至因引入误导性信息而降低性能。本文深入分析测试阶段的视觉想象作为可控资源的作用。首先研究静态视觉证据足够、想象能提升推理、过度或无意义想象影响准确率与效率的边界条件。为此,提出AVIC框架,通过显式判断当前视觉证据充分性,选择性地调用并控制想象强度。进一步提出AVIC-R,通过从问答正确性奖励与想象成本中学习策略,无需人工标注即可训练门控与规划行为。在空间推理基准(SAT、MMSI)和具身导航基准(R2R)上,结果揭示了想象关键、边际或有害的具体场景,并表明选择性控制可达到或超越固定策略,同时减少超过60%的世界模型调用和语言令牌消耗。AVIC-R优于GPT-4o和GPT-4.1等强基线,且调用频率更低。整体表明,分析与控制测试阶段的想象对高效可靠的视觉空间推理至关重要。

原文摘要 · Abstract (English)

Despite rapid progress in MLLMs, visual spatial reasoning remains unreliable when correct answers depend on how a scene would appear under unseen or alternative viewpoints. Recent work addresses this by augmenting reasoning with world models for visual imagination, but questions such as when imagination is actually necessary, how much of it is beneficial, and when it becomes harmful, remain poorly understood. In practice, indiscriminate imagination can increase computation and even degrade performance by introducing misleading evidence. In this work, we present an in-depth analysis of test-time visual imagination as a controllable resource for spatial reasoning. We first study when static visual evidence is sufficient, when imagination improves reasoning, and how excessive or unnecessary imagination affects accuracy and efficiency. To support this analysis, we then introduce AVIC, an adaptive test-time framework with world models that explicitly reasons about the sufficiency of current visual evidence before selectively invoking and scaling visual imagination. Finally, to further learn this gating and planning behavior without any annotation of when and how much to imagine, we introduce AVIC-R, which trains the policy via GRPO from QA-correctness rewards and penalties by imagination cost. Across spatial reasoning benchmarks (SAT, MMSI) and an embodied navigation benchmark (R2R), our results reveal clear scenarios where imagination is critical, marginal, or detrimental, and show that selective control can match or outperform fixed imagination strategies with substantially fewer world-model calls and language tokens. Our AVIC-R surpasses strong proprietary baselines including GPT-4o and GPT-4.1 while invoking the world model less often. Overall, our findings highlight the importance of analyzing and controlling test-time imagination for efficient and reliable spatial reasoning.

视觉推理世界模型自适应生成测试时优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。