arXiv:2605.10806cs.CVcs.AI2026-05被引 4

评测视频生成模型的物理推理能力,发现多数模型存在系统性错误。

PhyGround: Benchmarking Physical Reasoning in Generative World Models

论文配图:PhyGround: Benchmarking Physical Reasoning in Generative World Models
图 1 · 摘自论文原文
  • 设计13类物理定律的可测子问题,实现精准故障诊断。
  • 459名标注员提供超3.7万条细粒度标签,模型排名相关性超0.9。
  • 开源专用裁判模型PhyJudge-9B,偏差比Gemini低80%以上。

生成式世界模型在视频生成中日益重要,但评估其是否遵循真实物理规律仍具挑战。现有物理视频评测基准存在评估框架粗糙、标注偏见与疲劳、自动化评价器物理感知不足等问题。为此,我们提出PhyGround,一个基于准则的物理推理评测基准。包含250个精心设计的提示,每条附有预期物理结果,并构建涵盖固体力学、流体动力学和光学的13类物理定律分类体系。每类定律通过可观测子问题实现可量化诊断。我们通过大规模、质量控制的人工评测(459名标注员,5,796条完整标注,超37.4万条细粒度标签)评估8个主流视频生成模型。经质量筛选后,标注结果的分半模型排名相关性(Spearman's rho)>0.90。为支持可复现的自动化评估,我们发布PhyJudge-9B——一个开放的物理专精视觉语言模型裁判。PhyJudge-9B在聚合相对偏差上显著优于Gemini-3.1-Pro(3.3% vs. 16.6%)。所有数据、模型检查点与代码已公开于https://phyground.github.io/。

原文摘要 · Abstract (English)

Generative world models are increasingly used for video generation, where learned simulators are expected to capture the physical rules that govern real-world dynamics. However, evaluating whether generated videos actually follow these rules remains challenging. Existing physics-focused video benchmarks have made important progress, but they still face three key challenges, including the coarse evaluation frameworks that hide law-specific failures, response biases and fatigue that undermine the validity of annotation judgments, and automated evaluators that are insufficiently physics-aware or difficult to audit. To address those challenges, we introduce PhyGround, a criteria-grounded benchmark for evaluating physical reasoning in video generation. The benchmark contains 250 curated prompts, each augmented with an expected physical outcome, and a taxonomy of 13 physical laws across solid-body mechanics, fluid dynamics, and optics. Each law is operationalized through observable sub-questions to enable per-law diagnostics. We evaluate eight modern video generation models through a large-scale, quality-controlled human study, grounded on social science lab experiment design. A total of 459 annotators provided 5,796 complete annotations and over 37.4K fine-grained labels; after quality control, the retained annotations exhibited high split-half model-ranking correlations (Spearman's rho > 0.90). To support reproducible automated evaluation, we release PhyJudge-9B, an open physics-specialized VLM judge. PhyJudge-9B achieves substantially lower aggregate relative bias than Gemini-3.1-Pro (3.3% vs. 16.6%). We release prompts, human annotations, model checkpoints, and evaluation code on the project page https://phyground.github.io/.

视频生成物理推理评测基准VLM裁判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。