评测大模型对物理规律的理解,发现多数模型只靠外观猜测
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
- 设计真实与仿真环境结合的统一评测基准
- 模型在质心、杠杆平衡、牛顿第一定律上普遍出错
- 适合研究物理推理与世界模型的开发者参考
当前基础多模态大模型和视频世界模型在数学、常识和视觉推理方面进步显著,但对底层物理规律的理解仍不充分。现有评测多依赖合成图像或感知质量,难以衡量模型生成内容是否符合物理法则。为此,我们提出PhysicsMind,一个融合真实与仿真环境的统一基准,评估模型在质心、杠杆平衡和牛顿第一定律三个经典原理上的规律性推理与生成能力。包含两类任务:一为图像/短视频问答(VQA),检验模型能否从视觉输入中推断物理量;二为视频生成(VG)任务,评估预测轨迹是否满足真实运动中的质心、力矩与惯性约束。对多种前沿模型的测试表明,它们普遍依赖外观启发式,常违反基本力学规律。这说明当前规模与训练仍不足以实现稳健的物理理解,PhysicsMind可作为物理感知多模态模型的专用评测平台。数据将在论文接受后公开。
原文摘要 · Abstract (English)
Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp of the underlying physics remains underexplored. Existing benchmarks attempting to measure this matter rely on synthetic, Visual Question Answer templates or focus on perceptual video quality that is tangential to measuring how well the video abides by physical laws. To address this fragmentation, we introduce PhysicsMind, a unified benchmark with both real and simulation environments that evaluates law-consistent reasoning and generation over three canonical principles: Center of Mass, Lever Equilibrium, and Newton's First Law. PhysicsMind comprises two main tasks: i) VQA tasks, testing whether models can reason and determine physical quantities and values from images or short videos, and ii) Video Generation(VG) tasks, evaluating if predicted motion trajectories obey the same center-of-mass, torque, and inertial constraints as the ground truth. A broad range of recent models and video generation models is evaluated on PhysicsMind and found to rely on appearance heuristics while often violating basic mechanics. These gaps indicate that current scaling and training are still insufficient for robust physical understanding, underscoring PhysicsMind as a focused testbed for physics-aware multimodal models. Our data will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。