提出OPTIMA框架,让自动驾驶车在复杂场景中更安全高效协作。
OPTIMA: Optimized Policy for Intelligent Multi-Agent Systems Enables Coordination-Aware Autonomous Vehicles
- 通过交替采样与多智能体强化学习优化车辆协同决策。
- 在复杂拥挤场景下提升自动驾驶车的通用性与性能表现。
- 支持工业级分布式训练,适配多种算法与策略。
车联网与自动驾驶车辆(CAVs)之间的协调正因控制与通信技术的进步而发展。然而,当前多数研究基于过于简化且不现实的任务特定假设,可能引入潜在漏洞。由于CAVs不仅与环境互动,本身也是环境的一部分,若探索不足,政策可能隐含风险,因此亟需能广泛且高效探索环境的方法。本文提出OPTIMA,一种用于合作式自动驾驶任务的新型分布式强化学习框架。OPTIMA通过在环境交互中进行充分数据采样与多智能体强化学习算法间交替迭代,优化CAV间的协作,兼顾安全与效率。目标是提升CAVs在高度复杂、密集场景下的泛化能力与性能。此外,该工业级分布式训练系统可轻松适应不同算法、奖励函数与策略。
原文摘要 · Abstract (English)
Coordination among connected and autonomous vehicles (CAVs) is advancing due to developments in control and communication technologies. However, much of the current work is based on oversimplified and unrealistic task-specific assumptions, which may introduce vulnerabilities. This is critical because CAVs not only interact with their environment but are also integral parts of it. Insufficient exploration can result in policies that carry latent risks, highlighting the need for methods that explore the environment both extensively and efficiently. This work introduces OPTIMA, a novel distributed reinforcement learning framework for cooperative autonomous vehicle tasks. OPTIMA alternates between thorough data sampling from environmental interactions and multi-agent reinforcement learning algorithms to optimize CAV cooperation, emphasizing both safety and efficiency. Our goal is to improve the generality and performance of CAVs in highly complex and crowded scenarios. Furthermore, the industrial-scale distributed training system easily adapts to different algorithms, reward functions, and strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。