arXiv:2607.16401cs.CV2026-07被引 1

首个基于物理定律的视频生成模型评估基准,可诊断模型推理过程是否合理。

Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

论文配图:Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
图 1 · 摘自论文原文
  • 通过三阶段科学推理流程,从视觉感知到公式推导逐层检验模型逻辑
  • 11个模型最佳得分仅0.473,显示当前视频模型缺乏可靠物理推理能力
  • 适用于研究物理智能、世界模型与视频生成的开发者和评估者

当前视频生成模型常被视为具备物理规律内化能力的潜在世界模型,但现有评测仅关注输出结果的物理合理性,未验证其推理过程是否符合物理定律。本文提出Apple-PI,首个明确以物理定律为基础的视频模型评估基准。包含三个部分:1)Orchard数据集,含400段视频,涵盖经典力学中十类典型任务,区分单定律任务用于无混淆诊断,多定律任务用于泛化能力探测;2)基准评测协议,基于科学推理的三阶段流程——感知、建模、推导,采用信息图标注的首帧进行帧链提示,将生成视频视为模型的可视推理轨迹;3)混合评估套件,结合多模态大模型主观评分与基于物理定律的客观度量,实现对失败阶段、模块及来源的精准定位。对11个模型的评测表明,当前视频模型距离可靠的定律驱动世界模拟器仍很遥远,最优模型得分仅为0.473。阶段性分析揭示‘感知—建模—推导’瓶颈、多定律状态迁移弱以及持续存在的仿真到现实差距。Apple-PI为未来视频模型向具备物理智能的世界模型演进提供了诊断基础。

原文摘要 · Abstract (English)

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.

视频生成物理智能评测基准世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。