用拓扑结构提升自动驾驶车辆协同决策效率,突破传统方法的探索瓶颈。
Topology Enhanced MARL for Multi-Agent Cooperative Decision-Making of CAVs
- 将连续环境中的多智能体探索转化为有结构的拓扑遍历,降低维度灾难影响。
- 在真实测试中实现近最优决策分布,且具备零样本跨场景泛化能力。
- 适合自动驾驶协同控制、复杂交通场景下的强化学习研究者使用。
在连续环境中,去中心化的多智能体协同决策受维数灾难制约,盲目探索常收敛于保守局部最优。本文提出拓扑增强多智能体强化学习(TPE-MARL),将多智能体探索重构为结构化的拓扑遍历。引入博弈拓扑张量,利用局部敏感哈希将连续物理流形投影到离散商空间,该抽象作为对抗性信息瓶颈,解耦策略协调意图与环境噪声。在此空间中,双内在奖励机制驱动探索:新颖性奖励最大化已访问拓扑的边际熵,协作奖励通过变分证据下界(ELBO)优化以最小化条件熵,从而利用协同联合配置。实验表明,TPE-MARL达到接近蒙特卡洛树搜索(MCTS)预言器的理论最优决策分布。此外,物理测试平台验证了框架在零样本分布外(OOD)场景下的泛化能力。借助空间松弛机制,学习表征能可靠执行动态协商任务,如协作式交错并入,对真实世界协变量偏移和执行延迟具有内在鲁棒性。
原文摘要 · Abstract (English)
Decentralized multi-agent cooperative decision-making in continuous environments is fundamentally bottlenecked by the curse of dimensionality, where undirected exploration typically converges to conservative local optima. We propose Topology-Enhanced Multi-Agent Reinforcement Learning (TPE-MARL) to reformulate multi-agent exploration as a structured topological traversal. We introduce the Game Topology Tensor, utilizing locality-sensitive hashing to project the continuous physical manifold into a discrete quotient space. This abstraction operates as an adversarial Information Bottleneck, decoupling strategic coordination intents from environmental noise. Within this space, a dual intrinsic reward mechanism drives exploration: a novelty reward maximizes the marginal entropy of visited topologies, while a collaboration reward, optimized via a variational Evidence Lower Bound (ELBO), minimizes conditional entropy to exploit cooperative joint configurations. Evaluations demonstrate that TPE-MARL achieves near-optimal decision distributions, closely approximating the theoretical bounds established by a Monte Carlo Tree Search (MCTS) oracle. Furthermore, physical testbed experiments validate the framework's zero-shot out-of-distribution (OOD) generalization. Supported by a spatial relaxation mechanism, the learned representations reliably execute dynamic negotiations, such as cooperative zipper-merging, exhibiting inherent robustness against real-world covariate shifts and actuation latencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。