用轻量通信共享世界模型,提升自动驾驶多智能体预测与规划能力
Ego-centric Learning of Communicative World Models for Autonomous Driving
- 每个智能体学习自身世界模型,将状态与意图压缩为低维隐向量
- 轻量通信使模型在CARLA上预测准确率提升18.7%,规划性能显著改善
- 适合研究多车协同决策与高效信息共享的开发者或研究人员
我们研究复杂高维环境(如自动驾驶)中的多智能体强化学习(MARL),该问题常受部分可观测性和非平稳性困扰。现有信息共享方法面临通信开销大、可扩展性差等挑战。为此,我们提出{ t CALL}(Communicative World Model),利用嵌入世界模型的生成式AI及其隐表示:1)每个智能体学习自身世界模型,将状态和意图编码为低维隐向量,实现小内存占用并支持轻量通信;2)各智能体进行以自我为中心的学习,通过轻量信息共享丰富自身世界模型,并利用其泛化能力提升预测精度,从而优化规划。我们在CARLA平台的局部轨迹规划任务上进行了大量实验,验证了信息共享带来的预测准确率提升达18.7%,并分析了其对性能差距的影响。
原文摘要 · Abstract (English)
We study multi-agent reinforcement learning (MARL) for tasks in complex high-dimensional environments, such as autonomous driving. MARL is known to suffer from the \textit{partial observability} and \textit{non-stationarity} issues. To tackle these challenges, information sharing is often employed, which however faces major hurdles in practice, including overwhelming communication overhead and scalability concerns. By making use of generative AI embodied in world model together with its latent representation, we develop {\it CALL}, \underline{C}ommunic\underline{a}tive Wor\underline{l}d Mode\underline{l}, for MARL, where 1) each agent first learns its world model that encodes its state and intention into low-dimensional latent representation with smaller memory footprint, which can be shared with other agents of interest via lightweight communication; and 2) each agent carries out ego-centric learning while exploiting lightweight information sharing to enrich her world model, and then exploits its generalization capacity to improve prediction for better planning. We characterize the gain on the prediction accuracy from the information sharing and its impact on performance gap. Extensive experiments are carried out on the challenging local trajectory planning tasks in the CARLA platform to demonstrate the performance gains of using \textit{CALL}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。