首个量化评估视觉语言模型物理推理能力的基准,检验其对物体运动参数的数值判断能力。
QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- 构建3300+视频-文本对,要求模型估算物体尺寸、速度、加速度等数值属性。
- 现有顶尖模型在定性描述上看似合理,但实际数值误差显著,准确率不足40%。
- 揭示模型依赖预训练常识而非真实输入,适合关注模型可信赖性的研究者使用。
理解物理世界对通用人工智能代理至关重要。然而,当前最先进的视觉感知模型(如大视觉语言模型)是否具备定量物理推理能力仍不明确。现有评估多为基于问答的定性测试,难以判断模型能否从视频中推断运动物体的运动学量。为此,我们提出QuantiPhy,首个专门用于量化评估视觉语言模型物理推理能力的基准。该基准包含超过3.3K个带有数值真值的视频-文本实例,要求模型在给定某一属性作为先验的前提下,估计物体在特定时间戳的尺寸、速度和加速度。通过标准化提示与评分机制,可公平比较不同模型的数值准确性。实验发现,主流视觉语言模型在定性合理性与实际数值正确性之间存在明显差距。深入分析表明,模型严重依赖预训练的世界知识,而非忠实利用提供的视觉与文本输入进行定量推理。QuantiPhy为推动视觉语言模型从表面合理的语言表达迈向真正数值化的物理理解提供了首个严谨且可扩展的测试平台。
原文摘要 · Abstract (English)
Understanding the physical world is essential for generalist AI agents. However, it remains unclear whether state-of-the-art vision perception models (e.g., large VLMs) can reason physical properties quantitatively. Existing evaluations are predominantly VQA-based and qualitative, offering limited insight into whether these models can infer the kinematic quantities of moving objects from video observations. To address this, we present QuantiPhy, the first benchmark designed to quantitatively measure a VLM's physical reasoning ability. Comprising more than 3.3K video-text instances with numerical ground truth, QuantiPhy evaluates a VLM's performance on estimating an object's size, velocity, and acceleration at a given timestamp, using one of these properties as an input prior. The benchmark standardizes prompts and scoring to assess numerical accuracy, enabling fair comparisons across models. Our experiments on state-of-the-art VLMs reveal a consistent gap between their qualitative plausibility and actual numerical correctness. We further provide an in-depth analysis of key factors like background noise, counterfactual priors, and strategic prompting and find that state-of-the-art VLMs lean heavily on pre-trained world knowledge rather than faithfully using the provided visual and textual inputs as references when reasoning kinematic properties quantitatively. QuantiPhy offers the first rigorous, scalable testbed to move VLMs beyond mere verbal plausibility toward a numerically grounded physical understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。