多智能体强化学习让车辆安全并入高速车流,逼近理想控制效果。
A Systematic Study of Multi-Agent Deep Reinforcement Learning for Safe and Robust Autonomous Highway Ramp Entry
- 采用自对弈的多智能体深度强化学习,模拟复杂车流中多车协同变道。
- 在多车场景下实现近乎零碰撞,性能接近理论最优控制器。
- 适合研究自动驾驶协同决策与高阶自主系统设计的学者和工程师。
当前车辆可在高速公路上自动驾驶,无人出租车已在大城市运营,未来更高级别的自动驾驶将更加普及。然而,所谓的‘完全自动驾驶’(即L5级)尚未实现。要达成这一目标,必须具备如高速公路汇入车道等关键功能,并确保行为可证明安全且鲁棒可靠。本文系统研究了一种控制车辆纵向运动的高速公路汇入机制,旨在最小化汇入车辆与主路车流之间的碰撞风险。采用博弈论视角的多智能体(MA)方法,基于深度强化学习(DRL)构建控制器。虚拟环境通过自对弈生成仿真数据,使汇入车辆在渐缩型汇入过程中安全学习纵向位置控制策略。本文拓展了此前仅限两车交互的研究,系统性地增加交通流与汇入车辆数量。尽管已有研究表明,在完全去中心化、无协调的环境中,零碰撞控制器理论上不可行,但本文实证显示,所提方法训练出的控制器在性能上几乎达到理想最优水平。
原文摘要 · Abstract (English)
Vehicles today can drive themselves on highways and driverless robotaxis operate in major cities, with more sophisticated levels of autonomous driving expected to be available and become more common in the future. Yet, technically speaking, so-called "Level 5" (L5) operation, corresponding to full autonomy, has not been achieved. For that to happen, functions such as fully autonomous highway ramp entry must be available, and provide provably safe, and reliably robust behavior to enable full autonomy. We present a systematic study of a highway ramp function that controls the vehicles forward-moving actions to minimize collisions with the stream of highway traffic into which a merging (ego) vehicle enters. We take a game-theoretic multi-agent (MA) approach to this problem and study the use of controllers based on deep reinforcement learning (DRL). The virtual environment of the MA DRL uses self-play with simulated data where merging vehicles safely learn to control longitudinal position during a taper-type merge. The work presented in this paper extends existing work by studying the interaction of more than two vehicles (agents) and does so by systematically expanding the road scene with additional traffic and ego vehicles. While previous work on the two-vehicle setting established that collision-free controllers are theoretically impossible in fully decentralized, non-coordinated environments, we empirically show that controllers learned using our approach are nearly ideal when measured against idealized optimal controllers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。