arXiv:2506.09849cs.CV2025-06被引 40

测试深度模型对物理常识的理解,发现其表现远低于人类。

IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments

  • 基于四条物理原则设计虚拟场景测试模型判断事件是否合理
  • 多数模型在复杂场景中准确率仅50%,接近随机猜测
  • 适合研究通用智能与物理推理的学者参考

我们提出IntPhys 2,一个用于评估深度学习模型直观物理理解能力的视频基准。在原始IntPhys基础上,IntPhys 2聚焦宏观物体的四个核心原则:恒常性、不变性、时空连续性和固体性。这些原则源于早期儿童直觉物理认知的研究。IntPhys 2提供一套全面的测试,基于预期违背框架,挑战模型在受控且多样化的虚拟环境中区分可能与不可能事件的能力。同时,我们对多个顶尖模型进行了性能评估。结果表明,尽管这些模型具备基本视觉理解能力,但在复杂场景中对四大物理原则的直观理解仍面临巨大挑战,多数模型准确率仅为50%(即随机水平),与人类近乎完美的表现形成鲜明对比。这凸显了当前模型与人类级直觉物理理解之间的显著差距,强调了改进模型架构与训练方法的必要性。

原文摘要 · Abstract (English)

We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2 offers a comprehensive suite of tests, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments. Alongside the benchmark, we provide performance evaluations of several state-of-the-art models. Our findings indicate that while these models demonstrate basic visual understanding, they face significant challenges in grasping intuitive physics across the four principles in complex scenes, with most models performing at chance levels (50%), in stark contrast to human performance, which achieves near-perfect accuracy. This underscores the gap between current models and human-like intuitive physics understanding, highlighting the need for advancements in model architectures and training methodologies.

物理推理基准测试深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。