解决多智能体信念不一致下的协同决策问题,提升安全性和性能。
Towards Optimal Performance and Action Consistency Guarantees in Dec-POMDPs with Inconsistent Beliefs and Limited Communication
- 设计新框架,在信念不一致时仍能选择最优联合动作
- 提供动作一致性和性能的概率保证,通信仅在必要时触发
- 适合通信受限的自动驾驶、机器人协作等场景
在不确定性下的多智能体决策是实现自主系统有效与安全运行的基础。现实中,各智能体基于自身观测维护对环境的信念,并据此规划行动,但多数现有方法假设所有智能体在计划时拥有相同信念,即基于相同数据。这一假设在通信受限时难以成立。实际中,智能体常面临信念不一致的问题,导致协调失败、性能下降甚至安全隐患。本文提出一种新的去中心化框架,显式处理信念不一致,实现最优联合动作选择。该方法对开环多智能体部分可观马尔可夫决策过程(open-loop multi-agent POMDP)提供动作一致性与性能的概率保障,并仅在必要时触发通信。此外,还探讨了选定联合动作后是否应共享数据以提升推理性能。仿真结果表明,该方法优于当前先进算法。
原文摘要 · Abstract (English)
Multi-agent decision-making under uncertainty is fundamental for effective and safe autonomous operation. In many real-world scenarios, each agent maintains its own belief over the environment and must plan actions accordingly. However, most existing approaches assume that all agents have identical beliefs at planning time, implying these beliefs are conditioned on the same data. Such an assumption is often impractical due to limited communication. In reality, agents frequently operate with inconsistent beliefs, which can lead to poor coordination and suboptimal, potentially unsafe, performance. In this paper, we address this critical challenge by introducing a novel decentralized framework for optimal joint action selection that explicitly accounts for belief inconsistencies. Our approach provides probabilistic guarantees for both action consistency and performance with respect to open-loop multi-agent POMDP (which assumes all data is always communicated), and selectively triggers communication only when needed. Furthermore, we address another key aspect of whether, given a chosen joint action, the agents should share data to improve expected performance in inference. Simulation results show our approach outperforms state-of-the-art algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。