用双驱动强化学习提升自动驾驶在混行交通中的安全决策能力
Knowledge-Data-Dual-Driven Reinforcement Learning for Autonomous Vehicle Control in Mixed Traffic
- 融合生成模型与物理约束,主动预测车辆意图并生成前瞻性状态
- 在仿真中实现更安全、高效且舒适的驾驶表现,训练收敛更快
- 适合关注自动驾驶行为建模与多尺度控制的科研与工程人员
在混行交通中,自动驾驶车辆决策面临三大相互关联的挑战:首先,基于物理先验的强化学习模型难以捕捉周围车辆的潜在意图和多样化的驾驶行为,限制了主动推理能力;其次,周边车辆的突发动作导致非平稳性,使得长尾安全事件未被充分探索;第三,连续跟车与离散变道动作的混合动作空间引发统一强化学习训练的不稳定性。为此,本文提出知识-数据双驱动强化学习(KDDRL)。首先,条件深度生成模型合成意图感知的未来轨迹,将被动感知转化为主动预测状态;其次,知识-数据双驱动范式在这些预测状态上运行,融合概率性数据驱动洞察与物理约束,引导在安全关键场景下的安全探索;最后,耦合模块将意图感知轨迹与物理约束压缩为紧凑共享嵌入,实现连续跟车与离散变道的异步多时标优化,同时保留互信息。在数据集校准的仿真环境中评估表明,KDDRL有效应对意图不确定性,加速训练收敛,并在安全性、效率和舒适性上优于传统基线方法。
原文摘要 · Abstract (English)
In mixed traffic, decision-making for autonomous vehicles (AVs) confronts three interrelated challenges. First, physics-based priors incorporated into reinforcement learning (RL) models fail to capture latent interactive vehicle intentions and diverse driver behaviors, limiting the proactive reasoning capabilities. Second, abrupt maneuvers by surrounding vehicles cause non-stationarity, leaving long-tail safety events under-explored. Third, hybrid action spaces destabilize unified RL training due to the different temporal scales of continuous car-following and discrete lane-changing maneuvers. To address these issues, we propose Knowledge-Data Dual-driven Reinforcement Learning (KDDRL). First, a conditional deep generative model synthesizes intention-aware future trajectories, converting passive perception into proactive predictive states. Second, a knowledge-data dual-driven paradigm operates on these predictive states, fusing probabilistic data-driven insights with physical constraints to guide safe exploration through safety-critical scenarios. Third, a coupling module compresses both intention-aware trajectories and physical constraints into compact shared embeddings. This unified representation enables asynchronous multi-timescale optimization of continuous car-following and discrete lane-changing while preserving mutual information. Evaluations on dataset-calibrated simulations demonstrate that KDDRL effectively handles intention uncertainty, accelerates training convergence, and outperforms conventional baseline methods in terms of safety, efficiency, and comfort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。