arXiv:2502.20694cs.CVcs.AI2025-02NeurIPS被引 99

新基准评估视频生成模型对物理世界的建模能力,更贴近真实应用需求。

WorldModelBench: Judging Video Generation Models As World Models

  • 设计多维度评测框架,关注物体大小变化等物理规律违背。
  • 基于6.7万条人工标注,验证14个前沿模型在物理一致性上的表现。
  • 可自动评估且效果优于GPT-4o,适合研究视频世界模型的团队使用。

视频生成模型发展迅速,被视为能支持机器人、自动驾驶等决策应用的视频世界模型。然而现有评测基准仅关注视频整体质量,忽视物理一致性等世界模型关键因素。为此,我们提出WorldModelBench,一个面向应用领域的世界建模能力评测基准。该基准具备两大优势:(1) 检测细微世界建模违规行为,通过引入指令遵循与物理一致性维度,识别如物体尺寸异常变化违反质量守恒定律等问题,此类问题被以往基准忽略;(2) 与大规模人类偏好对齐,我们收集了67,000条人工标注,精准衡量14个前沿模型的表现。基于高质量人工标签,我们进一步微调出一个高精度评判器,实现自动化评估,预测世界建模违规的平均准确率比20亿参数的GPT-4o高出8.6%。此外,我们证明通过最大化评判器奖励进行训练,可显著提升模型的世界建模能力。项目官网:https://worldmodelbench-team.github.io。

原文摘要 · Abstract (English)

Video generation models have rapidly progressed, positioning themselves as video world models capable of supporting decision-making applications like robotics and autonomous driving. However, current benchmarks fail to rigorously evaluate these claims, focusing only on general video quality, ignoring important factors to world models such as physics adherence. To bridge this gap, we propose WorldModelBench, a benchmark designed to evaluate the world modeling capabilities of video generation models in application-driven domains. WorldModelBench offers two key advantages: (1) Against to nuanced world modeling violations: By incorporating instruction-following and physics-adherence dimensions, WorldModelBench detects subtle violations, such as irregular changes in object size that breach the mass conservation law - issues overlooked by prior benchmarks. (2) Aligned with large-scale human preferences: We crowd-source 67K human labels to accurately measure 14 frontier models. Using our high-quality human labels, we further fine-tune an accurate judger to automate the evaluation procedure, achieving 8.6% higher average accuracy in predicting world modeling violations than GPT-4o with 2B parameters. In addition, we demonstrate that training to align human annotations by maximizing the rewards from the judger noticeably improve the world modeling capability. The website is available at https://worldmodelbench-team.github.io.

视频生成世界模型评测基准物理一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。