用自进化大模型生成更真实且更具挑战性的自动驾驶测试场景
EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

- 基于大模型代理的自进化框架,模拟器约束下迭代优化场景生成
- 在MetaDrive和CARLA上显著扩展了安全与真实性的权衡边界
- 适合自动驾驶安全验证、强化学习训练及对抗性测试的研究者
生成安全性关键场景对于验证和改进自动驾驶系统至关重要,但需在最大化对抗性以暴露缺陷的同时保持现实性。现有方法通常依赖人工设计规则,受限于已知先验,忽略未探索模式。尽管近期开放式代理演化可突破此限制,但无约束通用代理缺乏模拟器约束,常将多目标冲突简化为单标量最大化。本文提出EvoDrive,首个基于大模型的自动化多目标场景生成框架。其采用模拟器接地的演员-评论家架构:记忆驱动的演员迭代提出生成器改进建议,评论家过滤不合理的候选,自演化世界评估器将有潜力的提议导向最优仿真预算。同时维护帕累托存档以保存多样化的攻击-现实权衡,并通过仿真反馈引导后续演化。在MetaDrive和CARLA上的基准测试表明,EvoDrive不仅显著扩展了各类生成器的帕累托前沿,还生成了对策略训练有价值的场景。
原文摘要 · Abstract (English)
Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversariality to expose failures while preserving realism. Existing methods usually manage this trade-off with handcrafted heuristics, confining generation to known priors and overlooking underexplored patterns. While recent open-ended agentic evolution can push this limit, unconstrained general agents lack strict simulator grounding and tend to collapse the multi-objective tension into single-scalar maximization. Here we present EvoDrive, the first automated, LLM-based agentic evolution framework for multi-objective scenario generation. EvoDrive employs a simulator-grounded actor-critic architecture where a memory-driven actor iteratively proposes improvements to the generators and critics filter out implausible candidates, and a self-evolving world evaluator routes promising proposals to optimize simulation budgets. EvoDrive further maintains a Pareto archive of evaluated candidates to preserve diverse attack-realism trade-offs and guide future evolution via simulation feedback. Benchmark results on MetaDrive and CARLA show that EvoDrive not only significantly expands the Pareto frontier across various generators, but also produces valuable scenarios for policy training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。