提出混合动作结构的强化学习方法,实现自动驾驶多目标兼容决策。
Hybrid Action Based Reinforcement Learning for Multi-Objective Compatible Autonomous Driving
- 采用多评价网络与混合参数化动作空间,解耦不同驾驶目标。
- 在仿真和HighD数据集上,效率、动作一致性与安全性均显著提升。
- 适合需要多目标平衡的复杂高速公路自动驾驶场景。
强化学习在自动驾驶决策与控制问题中表现优异,但驾驶是多属性任务,现有RL方法在策略更新与执行中难以兼顾多目标。一方面,单一价值评估网络限制了复杂场景下耦合目标的策略更新;另一方面,单一动作空间结构导致驾驶灵活性不足或行为波动大。为此,我们提出基于混合参数化动作的多目标集成评论家强化学习方法。该方法构建先进MORL架构,通过独立奖励函数使集成评论家聚焦不同目标;引入混合参数化动作空间,生成同时包含抽象引导与具体控制指令的动作;并设计基于不确定性的探索机制,支持混合动作快速学习多目标兼容策略。实验表明,在基于仿真和HighD数据集的多车道高速场景中,该方法在效率、动作一致性与安全性方面均表现出色。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has shown excellent performance in solving decision-making and control problems of autonomous driving, which is increasingly applied in diverse driving scenarios. However, driving is a multi-attribute problem, leading to challenges in achieving multi-objective compatibility for current RL methods, especially in both policy updating and policy execution. On the one hand, a single value evaluation network limits the policy updating in complex scenarios with coupled driving objectives. On the other hand, the common single-type action space structure limits driving flexibility or results in large behavior fluctuations during policy execution. To this end, we propose a Multi-objective Ensemble-Critic reinforcement learning method with Hybrid Parametrized Action for multi-objective compatible autonomous driving. Specifically, an advanced MORL architecture is constructed, in which the ensemble-critic focuses on different objectives through independent reward functions. The architecture integrates a hybrid parameterized action space structure, and the generated driving actions contain both abstract guidance that matches the hybrid road modality and concrete control commands. Additionally, an uncertainty-based exploration mechanism that supports hybrid actions is developed to learn multi-objective compatible policies more quickly. Experimental results demonstrate that, in both simulator-based and HighD dataset-based multi-lane highway scenarios, our method efficiently learns multi-objective compatible autonomous driving with respect to efficiency, action consistency, and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。