arXiv:2601.21282cs.CV2026-01

新基准WorldBench可分离评估世界模型对单一物理概念的理解能力。

WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts

  • 通过解耦物理概念,实现对单一物理规律的独立评估
  • 发现主流模型在物体持续性、摩擦系数等关键物理属性上表现不佳
  • 适合评估视频生成与世界模型的物理一致性,助力机器人训练

近年来,生成式基础模型(常称“世界模型”)在机器人规划和自主系统训练等关键任务中引发关注。为可靠部署,这些模型需具备高物理保真度,准确模拟真实世界动态。然而,现有基于物理的视频基准测试存在概念纠缠问题,单个测试同时评估多个物理定律,严重限制诊断能力。我们提出WorldBench,一种专为概念特异性、解耦评估设计的新视频基准,可严格隔离并评估单一物理概念或定律。为实现全面覆盖,我们设计两个层级的评测:1)高层次直观物理理解,如物体持续性、尺度/视角;2)低层次物理常数与材料属性,如摩擦系数、流体黏度,以精确衡量生成视频与现实的偏差。在WorldBench上评估顶尖视频世界模型时,我们发现其在特定物理概念上存在系统性失败,所有测试模型均缺乏生成真实世界交互所需的物理一致性。通过概念特异性评估,WorldBench提供更细致、可扩展的框架,用于严谨评估视频生成与世界模型的物理推理能力,为构建更鲁棒、通用的世界模型驱动学习铺平道路。

原文摘要 · Abstract (English)

Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, these models must exhibit high physical fidelity, accurately simulating real-world dynamics. Existing physics-based video benchmarks, however, suffer from entanglement, where a single test simultaneously evaluates multiple physical laws and concepts, fundamentally limiting their diagnostic capability. We introduce WorldBench, a novel video-based benchmark specifically designed for concept-specific, disentangled evaluation, allowing us to rigorously isolate and assess understanding of a single physical concept or law at a time. To make WorldBench comprehensive, we design benchmarks at two different levels: 1) an evaluation of intuitive physical understanding with higher level concepts such as object permanence or scale/perspective, and 2) an evaluation of low-level physical constants and material properties such as friction coefficients or fluid viscosity, allowing to measure excatly how far from reality generated videos are. When SOTA video-based world models are evaluated on WorldBench, we find specific patterns of failure in particular physics concepts, with all tested models lacking the physical consistency required to generate reliable real-world interactions. Through its concept-specific evaluation, WorldBench offers a more nuanced and scalable framework for rigorously evaluating the physical reasoning capabilities of video generation and world models, paving the way for more robust and generalizable world-model-driven learning.

世界模型物理理解视频生成基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。