通过状态转移梯度设计奖励,提升多车协同决策的训练效率与性能。
A Differentiated Reward Method for Reinforcement Learning based Multi-Vehicle Cooperative Decision-Making Algorithms
- 基于稳态转移系统分析交通流,将状态变化梯度融入奖励设计。
- 在不同自动驾驶渗透率下,训练收敛速度更快,交通效率提升12%以上。
- 适用于复杂交通场景,适合研究多智能体协同决策的学者与工程师。
强化学习(RL)在通过状态-动作-奖励反馈循环优化多车协同驾驶策略方面展现出巨大潜力,但仍面临样本效率低等挑战。本文提出一种基于稳态转移系统的差异化奖励方法,通过分析交通流特性,将状态转移梯度信息融入奖励设计,旨在优化多车协同决策中的动作选择与策略学习。该方法在MAPPO、MADQN和QMIX等算法中进行了验证,结果表明其显著加速了训练收敛,在不同自动驾驶渗透率下均优于中心化奖励等基线方法,在交通效率、安全性和动作合理性方面表现更优。此外,该方法展现出强可扩展性与环境适应性,为复杂交通场景下的多智能体协同决策提供了新思路。
原文摘要 · Abstract (English)
Reinforcement learning (RL) shows great potential for optimizing multi-vehicle cooperative driving strategies through the state-action-reward feedback loop, but it still faces challenges such as low sample efficiency. This paper proposes a differentiated reward method based on steady-state transition systems, which incorporates state transition gradient information into the reward design by analyzing traffic flow characteristics, aiming to optimize action selection and policy learning in multi-vehicle cooperative decision-making. The performance of the proposed method is validated in RL algorithms such as MAPPO, MADQN, and QMIX under varying autonomous vehicle penetration. The results show that the differentiated reward method significantly accelerates training convergence and outperforms centering reward and others in terms of traffic efficiency, safety, and action rationality. Additionally, the method demonstrates strong scalability and environmental adaptability, providing a novel approach for multi-agent cooperative decision-making in complex traffic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。