提出安全感知扩散框架,提升自动驾驶离线强化学习鲁棒性
DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving

- 分层扩散架构聚焦安全威胁,过滤环境冗余
- 复杂交通场景下成功率85%,平均奖励4295.68
- 适合高密度交通中需可靠决策的自动驾驶研究
尽管扩散模型能有效捕捉自动驾驶的多模态行为先验,但离线强化学习策略仍易受分布偏移、重尾风险信号、分布外(OOD)动作生成及高维状态冗余影响。为此,我们提出DiDrive,一种基于分布引导的离线扩散框架,包含两个协同组件:风险感知分层扩散(RHDif)架构与3DICE策略优化范式。在状态空间,RHDif通过低层级风险门控编码器与高层级上下文调制器,过滤环境冗余并聚焦于安全关键威胁。在动作空间,3DICE通过样本内校准引导、时空优化及集成候选排序,缓解分布外过估计与梯度震荡。在CARLA基准上的评估表明,相比IQL、CQL和Diffusion-QL等基线,DiDrive在包含60辆车辆的复杂高密度交通场景中表现更优,成功率达85%,平均奖励为4295.68,为安全自动驾驶决策提供了稳健路径。
原文摘要 · Abstract (English)
While diffusion models effectively capture multimodal behavioral priors for autonomous driving, offline reinforcement learning (RL) policies remain susceptible to distribution shift, heavy-tailed risk signals, out-of-distribution (OOD) action generation, and high-dimensional state redundancy. To address these challenges, we propose DiDrive, a distribution-guided offline diffusion framework featuring two synergistic components: the Risk-Aware Hierarchical Diffusion (RHDif) architecture and the 3DICE policy optimization paradigm. In the state space, RHDif utilizes a low-level risk-gated encoder and a high-level contextual modulator to filter environmental redundancy and focus on safety-critical threats. In the action space, 3DICE mitigates OOD overestimation and gradient oscillation through in-sample calibrated guidance, spatiotemporal optimization, and ensemble-based candidate ranking. Evaluations on the CARLA benchmark demonstrate DiDrive's superiority over baselines like IQL, CQL, and Diffusion-QL, particularly in complex, high-density traffic scenarios with 60 vehicles, where it achieves an 85% success rate and a 4295.68 average reward, providing a robust pathway for safe autonomous driving decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。