分层强化学习统一自动驾驶行为与控制,提升驾驶效率与安全
Multi-Timescale Hierarchical Reinforcement Learning for Unified Behavior and Control of Autonomous Driving
- 设计分时尺度的层级策略,高层输出长期驾驶指令,底层生成短期控制命令
- 在模拟器和HighD数据集上,驾驶效率、动作一致性与安全性均显著提升
- 支持结构化道路多模态行为建模,适合复杂高速多车道场景应用
强化学习在自动驾驶中应用日益广泛,但多数方法忽视策略结构设计。仅输出短时控制指令的策略会导致驾驶行为波动,而仅输出长时驾驶目标的策略难以实现行为与控制的统一优化。为此,本文提出一种多时尺度层次强化学习方法。该方法采用层次化策略结构,高层与低层强化学习策略联合训练,分别生成长时运动引导和短时控制指令。其中,运动引导通过混合动作显式表示,可捕捉结构化道路下的多模态驾驶行为,并支持增量式低层扩展状态更新。此外,设计了层次化安全机制保障多时尺度安全。在基于模拟器和HighD数据集的高速多车道场景评估中,本方法显著提升自动驾驶性能,有效提高驾驶效率、动作一致性和安全性。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) is increasingly used in autonomous driving (AD) and shows clear advantages. However, most RL-based AD methods overlook policy structure design. An RL policy that only outputs short-timescale vehicle control commands results in fluctuating driving behavior due to fluctuations in network outputs, while one that only outputs long-timescale driving goals cannot achieve unified optimality of driving behavior and control. Therefore, we propose a multi-timescale hierarchical reinforcement learning approach. Our approach adopts a hierarchical policy structure, where high- and low-level RL policies are unified-trained to produce long-timescale motion guidance and short-timescale control commands, respectively. Therein, motion guidance is explicitly represented by hybrid actions to capture multimodal driving behaviors on structured road and support incremental low-level extend-state updates. Additionally, a hierarchical safety mechanism is designed to ensure multi-timescale safety. Evaluation in simulator-based and HighD dataset-based highway multi-lane scenarios demonstrates that our approach significantly improves AD performance, effectively increasing driving efficiency, action consistency and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。