arXiv:2605.10858cs.CVcs.RO2026-05被引 3

提出新基准,评估自动驾驶世界模型的物理与行为真实性。

Is Your Driving World Model an All-Around Player?

论文配图:Is Your Driving World Model an All-Around Player?
图 1 · 摘自论文原文
  • 构建多维度评测体系,涵盖视觉、几何、闭环驾驶与人类感知。
  • 六款模型均无全优表现,最强者人类真实感仅得2-3分(满分10)。
  • 提供带理由的人类标注数据集和可解释的自动评估代理。

当前驾驶世界模型虽能生成逼真行车记录视频,但无一模型在所有方面都表现出色:部分模型纹理逼真却违背物理规律,另一些保持几何一致性却在闭环规划中失效。这一脱节暴露了关键缺陷——现有评估重在画面真实感,忽视行为合理性。本文提出WorldLens统一基准,从像素质量、4D几何到闭环驾驶与人类感知对齐,覆盖五个维度、24项标准指标。对六种代表性模型的评估显示,无一在全部维度领先:纹理丰富的模型违反几何规律,几何敏感模型缺乏行为一致性,最强模型的人类真实感评分仅为2-3分(满分10)。为弥合算法指标与人类感知差距,进一步构建WorldLens-26K数据集(含26,808条人类标注偏好及理由),并训练出基于这些判断的WorldLens-Agent视觉语言评估代理,实现可扩展、可解释的自动化评估。三者共同构成一个以物理与行为真实性为核心的生成世界评估生态系统。

原文摘要 · Abstract (English)

Today's driving world models can generate remarkably realistic dash-cam videos, yet no single model excels universally. Some generate photorealistic textures but violate basic physics; others maintain geometric consistency but fail when subjected to closed-loop planning. This disconnect exposes a critical gap: the field evaluates how real generated worlds appear, but rarely whether they behave realistically. We introduce WorldLens, a unified benchmark that measures world-model fidelity across the full spectrum, from pixel quality and 4D geometry to closed-loop driving and human perceptual alignment, through five complementary aspects and 24 standardized dimensions. Our evaluation of six representative models reveals that no existing approach dominates across all axes: texture-rich models violate geometry, geometry-aware models lack behavioral fidelity, and even the strongest performers achieve only 2-3 out of 10 on human realism ratings. To bridge algorithmic metrics with human perception, we further contribute WorldLens-26K, a 26,808-entry human-annotated preference dataset pairing numerical scores with textual rationales, and WorldLens-Agent, a vision-language evaluator distilled from these judgments that enables scalable, explainable auto-assessment. Together, the benchmark, dataset, and agent form a unified ecosystem for assessing generated worlds not merely by visual appeal, but by physical and behavioral fidelity.

世界模型自动驾驶评估基准人类感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。