测试时验证框架能提升视觉语言模型的空间推理能力
Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling
- 提出基于可验证空间断言的新型验证机制
- 在SAT-Real上显著提升推理准确率,减少动作偏差
- 揭示当前世界模型存在信息瓶颈,难以支持精细推理
视觉语言模型在需要多视角理解与具身视角转换的空间推理任务中仍受限。近期方法如MindJourney通过测试时扩展,利用世界模型生成动作条件轨迹,并由启发式验证器筛选有用视图。本文系统评估此类验证器在多个基准上的表现,发现其校准能力弱,随机评分常能达到相似答案熵降低效果,暴露出系统性动作偏差和不可靠奖励信号。为此,我们提出基于空间断言的验证框架(ViSA),将测试时奖励锚定于可验证的帧级微观断言。该框架在SAT-Real上持续提升空间推理性能,并通过更均衡探索行为纠正轨迹选择偏差。然而在更具挑战性的MMSI-Bench上,所有验证器(包括我们的)均未实现稳定缩放,表明当前世界模型构成信息瓶颈,想象视图无法增强细粒度推理。这些发现揭示了基于世界模型测试时验证的优劣与局限。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) remain limited in spatial reasoning tasks that require multi-view understanding and embodied perspective shifts. Recent approaches such as MindJourney attempt to mitigate this gap through test-time scaling where a world model imagines action-conditioned trajectories and a heuristic verifier selects helpful views from such trajectories. In this work, we systematically examine how such test-time verifiers behave across benchmarks, uncovering both their promise and their pitfalls. Our uncertainty-based analyses show that MindJourney's verifier provides little meaningful calibration, and that random scoring often reduces answer entropy equally well, thus exposing systematic action biases and unreliable reward signals. To mitigate these, we introduce a Verification through Spatial Assertions (ViSA) framework that grounds the test-time reward in verifiable, frame-anchored micro-claims. This principled verifier consistently improves spatial reasoning on the SAT-Real benchmark and corrects trajectory-selection biases through more balanced exploratory behavior. However, on the challenging MMSI-Bench, none of the verifiers, including ours, achieve consistent scaling, suggesting that the current world models form an information bottleneck where imagined views fail to enrich fine-grained reasoning. Together, these findings chart the bad, good, and ugly aspects of test-time verification for world-model-based reasoning. Our code is available at https://github.com/chandar-lab/visa-for-mindjourney.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。