arXiv:2606.03986cs.CV2026-06被引 1

用真实物理场景构建4D数据集,测试大模型对牛顿力学的理解能力。

NewtPhys: Do Foundation Models Understand Newtonian Physics?

论文配图:NewtPhys: Do Foundation Models Understand Newtonian Physics?
图 1 · 摘自论文原文
  • 基于多视角真实场景与物理模拟,生成带密集时序标注的4D数据集
  • 评估56个视觉语言模型在低层次物理推理上的表现,发现普遍局限
  • 适合研究物理感知视觉、下一代物理评测的学者使用

以往研究通过合成或半合成场景和视觉问答任务评估基础模型的物理推理能力,但这些基准侧重高层事件,缺乏评估底层牛顿物理理解所需的视觉保真度。我们提出NewtPhys,一个从真实世界场景的多视角图像构建的4D物理标注数据集,结合物理驱动的模拟。该数据集提供跨时间步的密集细粒度标注,包括3D受力、无模态像素级物理量、追踪、语义与几何信息,弥合了简化合成设置与真实视觉复杂性之间的差距。利用NewtPhys,我们系统评估了56个视觉语言模型(含54个开源模型和2个闭源前沿模型)及10个视觉-物理模型,揭示其在低层次物理推理中的局限性。除基准测试外,该数据集还推动物理引导视觉研究与下一代物理感知评估的发展。代码与数据集见https://astra-vision.github.io/NewtPhys。

原文摘要 · Abstract (English)

Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However, these benchmarks emphasize high-level events and lack the visual fidelity required to assess true low-level Newtonian understanding. We introduce NewtPhys, a 4D physically annotated dataset built from multiview images of real-world scenes with physics-grounded simulations. The dataset provides dense, fine-grained annotations across timesteps -- including 3D forces and amodal per-pixel quantities covering physics, tracking, semantics and geometry -- bridging the gap between simplistic synthetic setups and realistic visual complexity. Using NewtPhys, we systematically evaluate 56 VLMs, including 54 open-weight models and 2 closed-source frontier models, and 10 VFMs and reveal limitations in low-level physics reasoning. Beyond benchmarking, our dataset enables future research in physics-grounded vision and the development of next-generation physics-aware evaluations. Code and datasets are available at https://astra-vision.github.io/NewtPhys.

物理理解视觉语言模型4D数据集牛顿力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。