arXiv:2502.03960cs.RO2025-02中稿 · IEEE Transactions …被引 9

用双层强化学习让自动驾驶车在无红绿灯路口更懂周围车辆意图。

Bilevel Multi-Armed Bandit-Based Hierarchical Reinforcement Learning for Interaction-Aware Self-Driving at Unsignalized Intersections

  • 分层架构:高层用多臂老虎机应对车辆行为不确定性,指导底层控制。
  • 动态课程训练使样本效率提升,在CARLA中性能优于所有基线方法。
  • 适合复杂交互场景下的自动驾驶决策,尤其擅长处理未知车辆数量与意图。

本文提出BiM-ACPPO,一种基于双层多臂老虎机的分层强化学习框架,用于无红绿灯路口的交互感知决策与规划。该方法主动考虑周边车辆(SVs)带来的不确定性,包括驾驶意图、交互行为及数量变化。引入中间决策变量,使高层强化学习策略能生成交互感知参考,指导底层模型预测控制(MPC),从而增强框架泛化能力。利用无信号路口的结构特性,将强化学习训练建模为双层课程学习任务,通过提出的Exp3.S-based BiMAB算法求解。训练课程可动态调整,显著提升样本效率。在高保真CARLA仿真器中进行对比实验,结果表明本方法性能全面优于各基线。此外,在两个新城市驾驶场景中的实验也充分验证了其出色的泛化能力。

原文摘要 · Abstract (English)

In this work, we present BiM-ACPPO, a bilevel multi-armed bandit-based hierarchical reinforcement learning framework for interaction-aware decision-making and planning at unsignalized intersections. Essentially, it proactively takes the uncertainties associated with surrounding vehicles (SVs) into consideration, which encompass those stemming from the driver's intention, interactive behaviors, and the varying number of SVs. Intermediate decision variables are introduced to enable the high-level RL policy to provide an interaction-aware reference, for guiding low-level model predictive control (MPC) and further enhancing the generalization ability of the proposed framework. By leveraging the structured nature of self-driving at unsignalized intersections, the training problem of the RL policy is modeled as a bilevel curriculum learning task, which is addressed by the proposed Exp3.S-based BiMAB algorithm. It is noteworthy that the training curricula are dynamically adjusted, thereby facilitating the sample efficiency of the RL training process. Comparative experiments are conducted in the high-fidelity CARLA simulator, and the results indicate that our approach achieves superior performance compared to all baseline methods. Furthermore, experimental results in two new urban driving scenarios clearly demonstrate the commendable generalization performance of the proposed method.

自动驾驶强化学习交互感知决策规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。