提出新框架解决多智能体多目标决策中偏好不一致问题
Achieving Equilibrium under Utility Heterogeneity: An Agent-Attention Framework for Multi-Agent Multi-Objective Reinforcement Learning
- 用注意力机制隐式学习其他智能体的偏好和策略
- 在自定义环境和标准基准上显著优于现有方法
- 适合研究多智能体协作与冲突协调的学者
多智能体多目标系统(MAMOS)在机器人探索、自动驾驶交通管理、传感器网络优化等场景中展现出强大建模能力,通过去中心化控制提升可扩展性与鲁棒性,并更准确反映目标间的权衡。每个智能体使用效用函数将回报向量映射为标量值。现有方法难以处理效用函数异质性问题,因私有效用函数导致训练非平稳性加剧。本文首次理论证明:在去中心化执行约束下,实现贝叶斯纳什均衡需直接获取或结构化建模全局效用函数。为此,提出代理注意力多智能体多目标强化学习(AA-MAMORL)框架,在集中训练中隐式学习其他智能体效用函数及策略的联合信念,将全局状态与效用映射至各智能体策略。执行时,各智能体仅依赖局部观测与自身效用函数独立决策,无需通信即可近似贝叶斯纳什均衡。在自定义的MAMO Particle环境与标准MOMALand基准上的实验表明,获取全局偏好并采用本框架能显著提升性能,持续优于当前最优方法。
原文摘要 · Abstract (English)
Multi-agent multi-objective systems (MAMOS) have emerged as powerful frameworks for modelling complex decision-making problems across various real-world domains, such as robotic exploration, autonomous traffic management, and sensor network optimisation. MAMOS offers enhanced scalability and robustness through decentralised control and more accurately reflects inherent trade-offs between conflicting objectives. In MAMOS, each agent uses utility functions that map return vectors to scalar values. Existing MAMOS optimisation methods face challenges in handling heterogeneous objective and utility function settings, where training non-stationarity is intensified due to private utility functions and the associated policies. In this paper, we first theoretically prove that direct access to, or structured modeling of, global utility functions is necessary for the Bayesian Nash Equilibrium under decentralised execution constraints. To access the global utility functions while preserving the decentralised execution, we propose an Agent-Attention Multi-Agent Multi-Objective Reinforcement Learning (AA-MAMORL) framework. Our approach implicitly learns a joint belief over other agents' utility functions and their associated policies during centralised training, effectively mapping global states and utilities to each agent's policy. In execution, each agent independently selects actions based on local observations and its private utility function to approximate a BNE, without relying on inter-agent communication. We conduct comprehensive experiments in both a custom-designed MAMO Particle environment and the standard MOMALand benchmark. The results demonstrate that access to global preferences and our proposed AA-MAMORL significantly improve performance and consistently outperform state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。