arXiv:2508.10770cs.CV2025-08被引 3

诊断视觉语言模型的物理推理短板并有效提升其能力

From Diagnosis to Improvement: Probing Spatio-Physical Reasoning in Vision Language Models

  • 通过诊断发现模型依赖人类先验,缺乏深层推理
  • 微调加规则强化学习使模型超越主流闭源模型
  • 适合研究通用世界模型与多模态推理的学者

空间物理推理是理解真实物理世界的基础能力,对构建稳健的世界模型至关重要。尽管近期视觉语言模型(VLMs)在多模态数学和纯空间理解等特定领域取得显著进展,其在空间物理推理方面的能力仍基本未被探索。本文对主流VLMs进行了全面诊断分析,揭示当前模型在此关键任务上表现不佳。进一步分析表明,这种不足主要源于人类先验带来的偏差以及深层推理能力的缺失。为此,我们采用监督微调结合基于规则的强化学习对Qwen2.5-VL-7B进行优化,显著提升了其空间物理推理能力,并超越了领先的闭源模型。然而,该模型在新物理场景下的泛化能力仍然有限,凸显了空间物理推理领域亟需新方法。

原文摘要 · Abstract (English)

Spatio-physical reasoning, a foundation capability for understanding the real physics world, is a critical step towards building robust world models. While recent vision language models (VLMs) have shown remarkable progress in specialized domains like multimodal mathematics and pure spatial understanding, their capability for spatio-physical reasoning remains largely unexplored. This paper provides a comprehensive diagnostic analysis of mainstream VLMs, revealing that current models perform inadequately on this crucial task. Further detailed analysis shows that this underperformance is largely attributable to biases caused by human-like prior and a lack of deep reasoning. To address these challenges, we apply supervised fine-tuning followed by rule-based reinforcement learning to Qwen2.5-VL-7B, resulting in significant improvements in spatio-physical reasoning capabilities and surpassing leading proprietary models. Nevertheless, despite this success, the model's generalization to new physics scenarios remains limited -- underscoring the pressing need for new approaches in spatio-physical reasoning.

视觉语言模型物理推理模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。