评估世界模型在自动驾驶规划与因果关系中的表现,发现现有指标有局限性。
Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving

- 提出新评测框架,检验世界模型作为伪环境时对策略训练的可靠性。
- 实测显示顶尖模型在无扰动下表现好,但重放原始轨迹时大量失败。
- 新增敏感度指标,适合评估世界模型在因果场景下的规划能力。
世界模型作为学习型交通模拟器日益流行,已有研究尝试用其替代传统模拟器进行策略训练。本文探究现有评价指标是否适用于评估世界模型作为策略训练伪环境的性能。我们分析了Waymo Open Sim-Agents Challenge(WOSAC)所采用的元指标,并在标准场景中对比世界模型在完全或部分受控代理下的预测表现(部分回放)。此外,由于关注的是以自车行为条件化的世界模型,我们将标准WOSAC评估域扩展至包含对自车有因果影响的代理。实验发现,许多场景中排名靠前的模型在无扰动下表现良好,但在强制自车重放原始轨迹时出现严重失效。为此,我们提出了新的评价指标,用于揭示世界模型对不可控对象的敏感性,并评估其作为伪环境的性能。我们还分析了几种先进世界模型在新指标下的表现。
原文摘要 · Abstract (English)
World models have become increasingly popular in acting as learned traffic simulators. Recent work has explored replacing traditional traffic simulators with world models for policy training. In this work, we explore the robustness of existing metrics to evaluate world models as traffic simulators to see if the same metrics are suitable for evaluating a world model as a pseudo-environment for policy training. Specifically, we analyze the metametric employed by the Waymo Open Sim-Agents Challenge (WOSAC) and compare world model predictions on standard scenarios where the agents are fully or partially controlled by the world model (partial replay). Furthermore, since we are interested in evaluating the ego action-conditioned world model, we extend the standard WOSAC evaluation domain to include agents that are causal to the ego vehicle. Our evaluations reveal a significant number of scenarios where top-ranking models perform well under no perturbation but fail when the ego agent is forced to replay the original trajectory. To address these cases, we propose new metrics to highlight the sensitivity of world models to uncontrollable objects and evaluate the performance of world models as pseudo-environments for policy training and analyze some state-of-the-art world models under these new metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。