构建物理推理评测框架,揭示视觉语言模型在物理理解上的优劣。
Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models
- 设计400+问题的2D物理测试集,覆盖四大核心领域。
- 最大模型达0.815得分,规模越大推理能力越强。
- 擅长公式题,但空间抽象推理仍存短板,适合研究者参考。
随着视觉语言模型(VLMs)日益复杂,其推理能力受到越来越多关注。尽管在多项任务中表现优异,它们对基础科学原理(如物理)的理解仍处于探索阶段。为此,我们提出一个新颖且易用的评估框架,用于严格检验VLMs在二维物理理解方面的能力。该框架包含一个实用的问题生成器,构建了涵盖四个核心领域的超过400个测试问题:抛体运动、碰撞动力学、力学和流体动力学。通过对四个前沿VLMs的全面评估,我们发现模型规模与推理能力存在显著正相关,其中表现最佳的Qwen2.5-VL-7B模型总体得分为0.815。结果显示,模型在公式类问题上表现良好,但在需要抽象空间推理的领域仍面临显著挑战。通过该框架,我们旨在推动科学推理研究的普及,促进对VLMs能力与局限性的深入理解。
原文摘要 · Abstract (English)
As Vision-Language Models (VLMs) grow in sophistication, their ability to perform reasoning is coming under increasing supervision. While they excel at many tasks, their grasp of fundamental scientific principles, such as physics, remains an underexplored frontier. To reflect the advancements in these capabilities, we introduce a novel and accessible framework designed to rigorously evaluate VLMs on their understanding of 2D physics. Our framework features a pragmatic scenario generator that creates a diverse testbed of over 400 problems across four core domains: Projectile Motion, Collision Dynamics, Mechanics, and Fluid Dynamics. Through comprehensive evaluation of four state-of-the-art VLMs, we demonstrate a strong correlation between model scale and reasoning ability, with our top-performing model, Qwen2.5-VL-7B, achieving an overall score of 0.815. We find that while models excel at formulaic problems, they struggle significantly with domains requiring abstract spatial reasoning. By designing this framework, we aim to democratize the study of scientific reasoning in VLMs and foster deeper insights into their capabilities and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。