提出新方法诊断视觉语言模型在物理推理中的可信度问题。
NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics

- 构建双诊断框架,拆解物理推理中的视觉、物理规律与时间定位三要素。
- 6个主流VLM模型均因无法识别视觉前提或运用物理定律而表现不佳。
- 提供可校准置信度的新方法,适合关注模型可靠性与物理理解的研究者。
视觉语言模型(VLMs)在推导精确空间与物理洞察方面至关重要,但在空间智能任务如物理推理中表现不佳,成为主要瓶颈。当前缺乏科学分析以揭示模型是真实推理还是合理猜测。本文旨在深入理解VLMs如何感知物理世界及运用物理定律,并评估其置信度可靠性。提出NICE与FACT双诊断范式:FACT用于诊断视觉保真度、物理规律理解与时间定位;NICE则引入新型邻域感知校准方法与评估指标,以检验并校准置信度。在6个最新SOTA VLM上测试发现,模型普遍无法识别视觉前提或有效使用必要物理定律。本工作揭示了问题本质,并建立标准化诊断范式,推动更可信、物理根基扎实的VLM发展。
原文摘要 · Abstract (English)
The ability to derive precise spatial and physical insights is a cornerstone of vision-language models (VLMs), yet their poor performances in related spatial intelligence tasks such as physical reasoning remain a fundamental barrier. The community critically lacks a scientific analysis revealing whether VLMs faithfully reach answers or plausibly make guesses. This work aims to provide a fundamental understanding of how VLMs perceive the physical world, and utilize physical laws, while assessing the reliability of model confidence. We propose NICE and FACT, a dual-diagnostic paradigm that explicitly decomposes quantitative reasoning for kinematic physics: FACT diagnoses visual fidelity, physical law comprehension, and temporal grounding. NICE studies our novel neighborhood-informed calibration method and novel metrics to evaluate and calibrate confidence reliability. Evaluated across 6 latest state-of-the-art VLMs, we uncover that models fail to identify visual preconditions or utilize necessary physical laws to reach answers. This work highlights and establishes a standardized diagnostic paradigm to guide the development of faithful, physically-grounded VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。