arXiv:2607.10630cs.ROcs.AI2026-07

用自博弈对抗训练让自动驾驶规划更抗极端场景。

World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning

论文配图:World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning
图 1 · 摘自论文原文
  • 规划器与自生成对手在虚拟环境中互搏,动态学习危险应对策略。
  • 在nuPlan和InterPlan上实现92%的长尾场景成功率,优于基线方法。
  • 适合关注自动驾驶安全、规划鲁棒性的研究人员和工程师。

密集交通中的鲁棒运动规划需要自动驾驶车辆应对自然驾驶数据中罕见且高危的交互场景。尽管对抗训练可行,但现有方法依赖外部场景生成器、启发式扰动或高开销模拟器,难以与现代自回归规划器结合。本文将鲁棒规划器学习建模为约束极小极大博弈,提出对抗世界建模(AWM)框架。由于精确求解不可行,AWM采用理论合理的解耦求解器:内层最小化中,规划器的世界模型转化为角色条件对抗者,通过反事实信用分配学习稀疏、场景自适应的攻击联盟;外层最大化中,本车规划器优化带遗憾感知的鲁棒最优响应,利用尾部风险加权和参考锚定信任域,在保持正常驾驶行为的同时增强极端情况恢复能力。在nuPlan和InterPlan基准上的实验表明,该方法生成可迁移的对抗交互,闭环表现优异,涵盖正常与高度交互的长尾场景。理论分析验证了解耦求解器及核心优化组件的有效性。

原文摘要 · Abstract (English)

Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data. Although adversarial training offers a feasible solution, existing methods often rely on external scenario generators, heuristic perturbations, or simulator-heavy rollouts, which makes them difficult to integrate with modern autoregressive planners. Here, we cast adversarially robust planner learning as a constrained min-max game and propose Adversarial World Modeling (AWM), a theoretically grounded multi-agent self-play fine-tuning framework. Since solving the exact game is intractable, AWM introduces a principled decoupled solver. In the inner minimization, the planner's predictive world model is converted into a role-conditioned adversary that learns sparse, scene-adaptive attack coalitions via counterfactual credit assignment. In the outer maximization, the ego planner optimizes a regret-aware robust best response against the frozen AWM, utilizing tail-risk weighting and reference-anchored trust regions to improve hard-case recovery while preserving nominal driving behavior. Experiments on the nuPlan and InterPlan benchmarks demonstrate that our method generates transferable adversarial interactions and yields a robust planner that achieves competitive closed-loop performance in both nominal and highly interactive long-tail scenarios. Theoretical analysis justifies the decoupled solver and the main optimization components.

自动驾驶对抗训练运动规划自博弈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。