用感知结果指导自监督,解决自动驾驶闭环性能下降问题
Prioritizing Perception-Guided Self-Supervision: A New Paradigm for Causal Modeling in End-to-End Autonomous Driving
- 以感知输出为监督信号,显式建模环境与驾驶动作的因果关系
- 在Bench2Drive上取得78.08分驾驶得分和48.64%成功率,显著优于现有方法
- 适用于追求真实场景泛化能力的端到端自动驾驶系统研发
端到端自动驾驶系统主要通过模仿学习训练,虽在开环评估中表现良好,但在闭环场景下常因因果混淆导致性能大幅下降。这一问题根源在于模仿学习过度依赖包含不可归因噪声的专家轨迹,干扰了环境上下文与驾驶动作之间因果关系的建模。为此,我们提出感知引导的自监督(PGS)训练范式,利用感知输出(如车道中心线、周边车辆运动预测)作为主要监督信号,通过正负自监督对本车轨迹进行对齐,显式建模决策中的因果关系。该方法在标准端到端架构上实现78.08分的驾驶得分和48.64%的平均成功率,在挑战性的闭环Bench2Drive基准上显著优于现有最先进方法,包括使用更复杂网络结构和推理流程的方法。结果验证了PGS框架的有效性与鲁棒性,为解决因果混淆、提升自动驾驶真实场景泛化能力指明了新方向。
原文摘要 · Abstract (English)
End-to-end autonomous driving systems, predominantly trained through imitation learning, have demonstrated considerable effectiveness in leveraging large-scale expert driving data. Despite their success in open-loop evaluations, these systems often exhibit significant performance degradation in closed-loop scenarios due to causal confusion. This confusion is fundamentally exacerbated by the overreliance of the imitation learning paradigm on expert trajectories, which often contain unattributable noise and interfere with the modeling of causal relationships between environmental contexts and appropriate driving actions. To address this fundamental limitation, we propose Perception-Guided Self-Supervision (PGS) - a simple yet effective training paradigm that leverages perception outputs as the primary supervisory signals, explicitly modeling causal relationships in decision-making. The proposed framework aligns both the inputs and outputs of the decision-making module with perception results, such as lane centerlines and the predicted motions of surrounding agents, by introducing positive and negative self-supervision for the ego trajectory. This alignment is specifically designed to mitigate causal confusion arising from the inherent noise in expert trajectories. Equipped with perception-driven supervision, our method, built on a standard end-to-end architecture, achieves a Driving Score of 78.08 and a mean success rate of 48.64% on the challenging closed-loop Bench2Drive benchmark, significantly outperforming existing state-of-the-art methods, including those employing more complex network architectures and inference pipelines. These results underscore the effectiveness and robustness of the proposed PGS framework and point to a promising direction for addressing causal confusion and enhancing real-world generalization in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。