提出新基准,评估自动驾驶仿真中智能体的反应能力。
ReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving

- 分离车辆与周围智能体控制,强制智能体响应非日志行为。
- 构建2636个测试场景,验证多种模型在安全与规则合规性上的表现。
- 适合研究自动驾驶仿真、行为建模与反应能力评估的学者。
反应能力是数据驱动行为世界模型模拟器在自动驾驶仿真系统中的关键特性。具备该能力的模拟系统可使仿真智能体对与真实日志不同的自动驾驶车辆(AV)行为做出合理响应。然而,现有行为仿真基准未直接衡量反应能力,通常让模拟器同时控制AV和周边智能体,并通过日志相似性或开环预测指标评估真实性。本文提出ReactSim-Bench,用于评估自动驾驶行为世界模型的反应能力。我们解耦智能体与AV的控制,采用不同于日志的AV行为作为输入,要求智能体独立作出反应。为生成这些行为,我们构建了一个基于AV规划模型的候选行为生成管道,并通过规则过滤与人工验证筛选数据。使用碰撞率、地图相关性及运动学可行性等指标评估反应行为的安全性与规则符合性。我们构建了包含三类场景的2636个测试用例,对多种主流模型(基于Transformer、扩散模型、下一词预测模型)进行系统评估,并分析重规划频率对性能的影响,为未来研究提供洞见。
原文摘要 · Abstract (English)
Reactive capability is a key property of data-driven behavior world model simulators for autonomous driving simulation systems. With this capability, simulated world agents can respond feasibly to autonomous vehicle (AV) behaviors that differ from the log. However, existing behavior simulation benchmarks do not directly measure reactive capability. They often let the simulator jointly control the AV and surrounding agents and evaluate realism through log similarity or open-loop prediction metrics. In this work, we introduce ReactSim-Bench for evaluating the reactive capability of behavior world model simulation in autonomous driving. We decouple the control of agents and the AV, using AV behaviors that differ from the log and require agents to respond as independent AV inputs. To obtain these AV behaviors, we construct a pipeline that uses an AV planner model to generate candidate behaviors and filters the data using rules and manual verification. Collision metrics, map-based metrics, and kinematic feasibility metrics are used to evaluate the safety and rule compliance of reactive responses. We construct 2,636 test scenarios with three categories and conduct a systematic evaluation of state-of-the-art models across multiple architectures, including Transformer-based, diffusion-based, and next-token-prediction-based models. We further analyze how replan frequency affects performance and provide insights for future studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。