arXiv:2602.23259cs.CVcs.AI2026-02被引 2

不依赖专家数据,用风险感知模型实现更安全的自动驾驶决策。

Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving

  • 构建风险感知世界模型,预测多种动作后果并选择低风险行为。
  • 通过危险行为训练使模型可预判事故,避免极端风险。
  • 无需专家示范即可生成安全动作,适合长尾场景应对。

随着模仿学习(IL)和大规模驾驶数据集的发展,端到端自动驾驶(E2E-AD)取得显著进展。当前基于模仿学习的方法以专家行为为标准,通过最小化与专家动作的差异进行训练。然而,这种‘仅模仿专家’的目标在分布外场景下泛化能力有限:面对罕见或未见的长尾情况时,模型因缺乏先验经验而产生不安全决策。这引发一个根本问题:能否在无专家动作监督的情况下实现可靠的自动驾驶?为此,我们提出统一框架风险感知世界模型预测控制(RaWMPC),通过鲁棒控制解决泛化困境,且不依赖专家示范。实践中,RaWMPC利用世界模型预测多个候选动作的后果,并通过显式风险评估选择低风险动作。为使世界模型具备预测危险驾驶行为后果的能力,我们设计了风险感知交互策略,系统性地让模型暴露于高危行为,使灾难性结果可预测、可规避。此外,为在测试时生成低风险候选动作,我们引入自评估蒸馏方法,将训练好的世界模型中的避险能力蒸馏至生成式动作提案网络,全程无需专家示范。大量实验表明,RaWMPC在分布内与分布外场景中均优于现有方法,且具有更优的决策可解释性。

原文摘要 · Abstract (English)

With advances in imitation learning (IL) and large-scale driving datasets, end-to-end autonomous driving (E2E-AD) has made great progress recently. Currently, IL-based methods have become a mainstream paradigm: models rely on standard driving behaviors given by experts, and learn to minimize the discrepancy between their actions and expert actions. However, this objective of "only driving like the expert" suffers from limited generalization: when encountering rare or unseen long-tail scenarios outside the distribution of expert demonstrations, models tend to produce unsafe decisions in the absence of prior experience. This raises a fundamental question: Can an E2E-AD system make reliable decisions without any expert action supervision? Motivated by this, we propose a unified framework named Risk-aware World Model Predictive Control (RaWMPC) to address this generalization dilemma through robust control, without reliance on expert demonstrations. Practically, RaWMPC leverages a world model to predict the consequences of multiple candidate actions and selects low-risk actions through explicit risk evaluation. To endow the world model with the ability to predict the outcomes of risky driving behaviors, we design a risk-aware interaction strategy that systematically exposes the world model to hazardous behaviors, making catastrophic outcomes predictable and thus avoidable. Furthermore, to generate low-risk candidate actions at test time, we introduce a self-evaluation distillation method to distill riskavoidance capabilities from the well-trained world model into a generative action proposal network without any expert demonstration. Extensive experiments show that RaWMPC outperforms state-of-the-art methods in both in-distribution and out-of-distribution scenarios, while providing superior decision interpretability.

自动驾驶风险感知世界模型模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。