测试多模态大模型在台球物理推理中的表现,发现其越复杂越容易出错。
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs

- 用程序化生成的台球场景测试模型对运动和碰撞的预测能力
- 模型在长时间模拟和复杂几何下准确率显著下降
- 发现模型普遍存在'静止偏差',倾向于预测无互动
当前多模态模型在静态图像识别上表现良好,但在直观物理推理方面仍显不足。仅凭单张图像预测物体如何运动和相互作用仍是难题。我们提出 BilliardPhys-Bench,一个基于合成台球环境的物理推理基准。其程序化引擎可生成含摩擦与弹性碰撞的随机场景。该基准测试三种能力:(1) 球与球的碰撞预测,(2) 墙面反弹推理,(3) 运动停止后的最终球位估计。我们评估了 GPT、Claude、Gemini 与 Qwen 系列的最新多模态大模型(MLLMs)。结果表明,随着模拟时间增加和场景几何复杂度提升,性能持续下降。还观察到一种一致的失败模式——'静止偏差':当正确物理结果更难推断时,模型倾向于预测无交互。这些发现揭示了当前多模态大模型在视觉动态推理中的局限性,并指向需要更强的物理先验知识以改进模型架构。
原文摘要 · Abstract (English)
Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move and interact from a single image is still difficult for these systems. We present BilliardPhys-Bench, a benchmark for physical reasoning in synthetic billiards environments. Its procedural engine generates randomized scenarios with friction and elastic collisions. The benchmark tests three abilities: (1) predicting ball-to-ball collisions, (2) reasoning about wall bounces, and (3) estimating final ball positions after motion stops. We evaluate recent MLLMs from the GPT, Claude, Gemini, and Qwen families. Performance drops as simulation time increases and scene geometry grows more complex. We also observe a consistent failure mode we call "stasis bias": when the correct physical outcome is harder to infer, models tend to predict no interaction. These findings show where current MLLMs break down on visual dynamics and point toward the need for better physical inductive biases in multimodal architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。