诊断大模型物理推理失败原因,发现其答对题可能靠巧合而非真理解。
Physics Knowledge in Frontier Models: A Diagnostic Study of Failure Modes
- 设计细分测试项,分离感知与物理理解能力
- 发现模型答对题却未必掌握真实物理规律
- 适合关注模型可靠性与可解释性的研究者
尽管近期视觉-语言模型(VLMs)在复杂推理任务上取得显著进展,但难以判断其成功或失败的原因。传统基准仅评估答案正确率,无法揭示背后的机制。本文针对六种前沿VLMs,在三个基于物理的基准(Physion、Physion++ 和 CLEVRER)上开展故障模式分析,通过自定义子测试(Physion 和 Physion++)及现有类别整合(CLEVRER),将性能分解为可验证的能力维度。这些子测试分别隔离感知(物体、颜色、遮挡识别)和物理理解(运动预测、空间推理),从而检验模型是否基于正确实体与动态作出判断。出人意料的是,子测试掌握程度与总分相关性很弱:模型常在未真正理解感知或物理的前提下答对。这表明当前VLMs可能因错误理由获得高分,凸显了超越平均指标、揭露隐藏失败模式诊断方法的重要性。
原文摘要 · Abstract (English)
While recent Vision-Language Models (VLMs) have achieved impressive progress, it remains difficult to determine why they succeed or fail on complex reasoning tasks. Traditional benchmarks evaluate what models can answer correctly, not why they succeed or fail. In this work, we perform a failure-mode analysis of six frontier VLMs on three physics-based benchmarks - Physion, Physion++, and CLEVRER - by introducing custom subtests (for Physion and Physion++) and an integration of existing benchmark categories (for CLEVRER) to factor benchmark performance into distinct, testable capabilities. These subtests isolate perception (object, color, and occlusion recognition) and physics understanding (motion prediction and spatial reasoning), enabling us to test whether models attend to the correct entities and dynamics underlying their answers. Counterintuitively, subtest mastery correlates only weakly with benchmark accuracy: models often answer correctly without grounding in perception or physics. This suggests that current VLMs sometimes achieve benchmark scores for the wrong reasons, underscoring the need for diagnostics that expose hidden failure modes beyond aggregate metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。