arXiv:2412.05337cs.CVcs.LG2024-12被引 12

提出可控制动作的自动驾驶世界模型评估框架,解决仿真与指令不一致问题

ACT-Bench: Towards Action Controllable World Models for Autonomous Driving

  • 构建开放评测框架ACT-Bench,结合nuScenes视频与轨迹数据
  • 新模型Terra在动作指令遵循上显著优于现有顶尖模型
  • 适合自动驾驶仿真、具身智能研究者使用

世界模型作为自动驾驶的神经仿真器具有潜力,可弥补真实数据稀缺并支持闭环评估。然而当前研究多以视觉真实度或下游任务表现评价模型,对动作指令忠实度关注不足——这正是生成特定场景的关键。尽管部分研究涉及动作忠实度,但其评估依赖闭源机制,难以复现。为此,我们开发了开源评估框架ACT-Bench,并提出基准模型Terra。该框架包含大规模数据集,将nuScenes中的短时上下文视频与对应未来轨迹数据配对,为生成未来视频帧提供条件输入,从而量化动作忠实度。Terra基于多个大规模轨迹标注数据集训练,强化动作忠实性。实验表明,当前最先进模型未能完全遵循指令,而Terra显著提升动作忠实度。所有组件将公开,以支持后续研究。

原文摘要 · Abstract (English)

World models have emerged as promising neural simulators for autonomous driving, with the potential to supplement scarce real-world data and enable closed-loop evaluations. However, current research primarily evaluates these models based on visual realism or downstream task performance, with limited focus on fidelity to specific action instructions - a crucial property for generating targeted simulation scenes. Although some studies address action fidelity, their evaluations rely on closed-source mechanisms, limiting reproducibility. To address this gap, we develop an open-access evaluation framework, ACT-Bench, for quantifying action fidelity, along with a baseline world model, Terra. Our benchmarking framework includes a large-scale dataset pairing short context videos from nuScenes with corresponding future trajectory data, which provides conditional input for generating future video frames and enables evaluation of action fidelity for executed motions. Furthermore, Terra is trained on multiple large-scale trajectory-annotated datasets to enhance action fidelity. Leveraging this framework, we demonstrate that the state-of-the-art model does not fully adhere to given instructions, while Terra achieves improved action fidelity. All components of our benchmark framework will be made publicly available to support future research.

世界模型自动驾驶动作控制仿真评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。