arXiv:2509.17940cs.ROcs.CV2025-09NeurIPS被引 39

用安全偏好优化提升端到端自动驾驶的驾驶安全性。

DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving

  • 融合人类模仿与规则安全得分,统一构建策略分布进行优化。
  • 在NAVSIM上达到90.0的PDMS新基准,显著提升安全表现。
  • 适合关注自动驾驶安全与策略优化的研究者和工程师。

端到端自动驾驶通过直接从原始感知输入预测未来轨迹取得了显著进展,但主流的模仿学习方法存在严重安全缺陷,无法区分看似类人却潜在危险的轨迹。部分近期方法尝试通过回归多规则驱动评分来解决,但监督信号与策略优化解耦,导致性能不佳。为此,我们提出DriveDPO——一种基于安全直接偏好优化的策略学习框架。首先,将人类模仿相似性与规则驱动安全得分融合,生成统一策略分布以支持直接策略优化;其次,引入轨迹级偏好对齐的迭代直接偏好优化阶段。在NAVSIM基准上的大量实验表明,DriveDPO实现了90.0的最新状态性能(PDMS)。定性结果进一步验证了其在多样化复杂场景中生成更安全、更可靠的驾驶行为的能力。

原文摘要 · Abstract (English)

End-to-end autonomous driving has substantially progressed by directly predicting future trajectories from raw perception inputs, which bypasses traditional modular pipelines. However, mainstream methods trained via imitation learning suffer from critical safety limitations, as they fail to distinguish between trajectories that appear human-like but are potentially unsafe. Some recent approaches attempt to address this by regressing multiple rule-driven scores but decoupling supervision from policy optimization, resulting in suboptimal performance. To tackle these challenges, we propose DriveDPO, a Safety Direct Preference Optimization Policy Learning framework. First, we distill a unified policy distribution from human imitation similarity and rule-based safety scores for direct policy optimization. Further, we introduce an iterative Direct Preference Optimization stage formulated as trajectory-level preference alignment. Extensive experiments on the NAVSIM benchmark demonstrate that DriveDPO achieves a new state-of-the-art PDMS of 90.0. Furthermore, qualitative results across diverse challenging scenarios highlight DriveDPO's ability to produce safer and more reliable driving behaviors.

自动驾驶偏好优化安全强化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。