让多个智能体学会协调各自的偏好,更好处理多目标冲突。
Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning

- 设计新算法让每个智能体学独立偏好,实现互补权衡
- 在多个协作环境和交通控制场景中提升性能与协调性
- 适合研究多智能体协同决策或实际应用的开发者
协作式多目标多智能体强化学习(MOMARL)模型在多个可能冲突的目标下进行团队决策。此设定中,冲突不仅存在于目标之间,也存在于观测、角色和贡献不同的智能体之间。我们提出偏好协调多智能体策略优化(PCMA),通过学习各智能体特有的偏好,实现智能体间的互补权衡。理论上,我们将协作式MOMARL建模为团队最优博弈,并证明在适当条件下,偏好多样性可通过一阶改进分解促进团队整体提升。在多个协作式多智能体环境及一个实际交通控制场景中的实验表明,PCMA显著提升了性能与权衡协调能力。
原文摘要 · Abstract (English)
Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, potentially conflicting objectives. In this setting, conflicts arise not only across objectives but also across agents with different observations, roles, and contributions. We propose Preference Coordinated Multi-agent Policy Optimization (PCMA), which learns coordinated agent-specific preferences to enable complementary trade-offs among agents. Theoretically, we formulate cooperative MOMARL as a team-optimal game and show that, under suitable conditions, preference diversity can induce team improvement through a first-order improvement decomposition. Experiments on multiple cooperative MOMA environments and a practical traffic-control scenario show that PCMA improves both performance and trade-off coordination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。