arXiv:2511.21584cs.ROcs.AI2025-11NeurIPS被引 11

用模拟生成反事实轨迹,提升自动驾驶模型在真实闭环场景下的安全性和鲁棒性。

Model-Based Policy Adaptation for Closed-Loop End-to-End Autonomous Driving

  • 基于几何一致的仿真生成多种反事实驾驶轨迹,扩充训练数据
  • 通过扩散模型适配器和多步Q值模型优化决策,提升长期表现
  • 适用于需高安全性与强泛化能力的自动驾驶部署场景

端到端自动驾驶模型在开环评估中表现优异,但在闭环设置下常因误差累积和泛化能力差而失效。为此,我们提出基于模型的策略适配(MPA)框架,增强预训练端到端驾驶代理在部署中的鲁棒性与安全性。MPA首先利用几何一致性仿真引擎生成多样化的反事实轨迹,使代理接触原始数据集外的场景。基于生成数据,训练一个基于扩散模型的策略适配器以修正基础策略预测,并构建多步Q值模型评估长期结果。推理时,适配器生成多个轨迹候选,由Q值模型选择预期效用最高的方案。在nuScenes基准上使用逼真闭环仿真器的实验表明,MPA显著提升了模型在域内、域外及高危场景下的性能。我们还研究了反事实数据规模与推理时引导策略对整体效果的影响。

原文摘要 · Abstract (English)

End-to-end (E2E) autonomous driving models have demonstrated strong performance in open-loop evaluations but often suffer from cascading errors and poor generalization in closed-loop settings. To address this gap, we propose Model-based Policy Adaptation (MPA), a general framework that enhances the robustness and safety of pretrained E2E driving agents during deployment. MPA first generates diverse counterfactual trajectories using a geometry-consistent simulation engine, exposing the agent to scenarios beyond the original dataset. Based on this generated data, MPA trains a diffusion-based policy adapter to refine the base policy's predictions and a multi-step Q value model to evaluate long-term outcomes. At inference time, the adapter proposes multiple trajectory candidates, and the Q value model selects the one with the highest expected utility. Experiments on the nuScenes benchmark using a photorealistic closed-loop simulator demonstrate that MPA significantly improves performance across in-domain, out-of-domain, and safety-critical scenarios. We further investigate how the scale of counterfactual data and inference-time guidance strategies affect overall effectiveness.

自动驾驶策略适配闭环控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。