arXiv:2604.07378cs.RO2026-04

让自动驾驶测试自动进化,从静态评估变为动态训练闭环。

Evaluation as Evolution: Transforming Adversarial Diffusion into Closed-Loop Curricula for Autonomous Vehicles

  • 将对抗生成改为可迭代优化的演化课程,动态发现危险场景。
  • 在nuScenes和nuPlan上碰撞失败发现率提升9.01%至21.43%。
  • 适合追求高鲁棒性的自动驾驶系统研发团队使用。

自动驾驶车辆在交互式交通环境中常受限于静态数据集中的安全关键尾部事件稀缺,导致学习策略偏向平均情况,降低鲁棒性。现有评估方法虽尝试通过对抗压力测试缓解此问题,但多为开环且事后进行,难以将发现的缺陷反馈至训练过程。本文提出评价即进化(Evaluation as Evolution, $E^2$),构建一个闭环框架,将对抗场景生成从静态验证转变为自适应演化课程。具体地,$E^2$ 将对抗场景合成建模为对学习到的逆时SDE先验的运输正则化稀疏控制。为应对高维生成挑战,采用拓扑驱动的支持选择识别关键交互主体,并引入拓扑锚定以稳定生成过程。该方法可在严格约束偏离真实数据分布的前提下,精准发现失效案例。实验表明,$E^2$ 在nuScenes数据集上碰撞失败发现率提升9.01%,在nuPlan上最高提升21.43%,同时保持低无效性和高真实性。进一步地,将生成的边界案例用于闭环策略微调,显著提升模型鲁棒性。

原文摘要 · Abstract (English)

Autonomous vehicles in interactive traffic environments are often limited by the scarcity of safety-critical tail events in static datasets, which biases learned policies toward average-case behaviors and reduces robustness. Existing evaluation methods attempt to address this through adversarial stress testing, but are predominantly open-loop and post-hoc, making it difficult to incorporate discovered failures back into the training process. We introduce Evaluation as Evolution ($E^2$), a closed-loop framework that transforms adversarial generation from a static validation step into an adaptive evolutionary curriculum. Specifically, $E^2$ formulates adversarial scenario synthesis as transport-regularized sparse control over a learned reverse-time SDE prior. To make this high-dimensional generation tractable, we utilize topology-driven support selection to identify critical interacting agents, and introduce Topological Anchoring to stabilize the process. This approach enables the targeted discovery of failure cases while strictly constraining deviations from realistic data distributions. Empirically, $E^2$ improves collision failure discovery by 9.01% on the nuScenes dataset and up to 21.43% on the nuPlan dataset over the strongest baselines, while maintaining low invalidity and high realism. It further yields substantial robustness gains when the resulting boundary cases are recycled for closed-loop policy fine-tuning.

自动驾驶对抗生成闭环训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。