arXiv:2508.05405cs.AI2025-08AAAI被引 11

评测视觉语言模型在物理推理任务中的表现,发现顶尖模型仍难精准控制。

DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning

  • 构建模拟环境评估模型对物理规律的理解与推理能力
  • 顶尖模型在复杂场景中仍无法实现精准动作规划
  • 适合关注具身智能与物理推理的科研人员参考

尽管视觉语言模型(VLMs)具备强大的感知能力和出色的视觉推理能力,但在复杂动态环境中仍难以关注细节和进行精确的动作规划,导致性能不佳。现实任务通常需要复杂的交互、高级的空间推理、长期规划及持续策略优化,这往往依赖于对目标场景物理规则的理解。然而,在真实场景中评估这些能力通常成本过高。为此,我们提出DeepPHY,一个新颖的基准框架,通过一系列具有挑战性的模拟环境,系统性地评估VLMs对基础物理原理的理解与推理能力。DeepPHY整合了不同难度级别的物理推理环境,并引入细粒度评估指标。我们的评估发现,即使是最先进的VLMs也难以将描述性物理知识转化为精确的预测性控制。

原文摘要 · Abstract (English)

Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in complex, dynamic environments, leading to subpar performance. Real-world tasks typically require complex interactions, advanced spatial reasoning, long-term planning, and continuous strategy refinement, usually necessitating understanding the physics rules of the target scenario. However, evaluating these capabilities in real-world scenarios is often prohibitively expensive. To bridge this gap, we introduce DeepPHY, a novel benchmark framework designed to systematically evaluate VLMs' understanding and reasoning about fundamental physical principles through a series of challenging simulated environments. DeepPHY integrates multiple physical reasoning environments of varying difficulty levels and incorporates fine-grained evaluation metrics. Our evaluation finds that even state-of-the-art VLMs struggle to translate descriptive physical knowledge into precise, predictive control.

物理推理视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。