arXiv:2412.12326cs.MAcs.AI2024-12被引 4

通过共享行动建议,让智能体在利益冲突中实现集体福祉。

Achieving Collective Welfare in Multi-Agent Reinforcement Learning via Suggestion Sharing

  • 智能体间交换行动建议,而非共享奖励或策略。
  • 理论证明建议共享可缩小个体与集体目标的差距。
  • 无需设计内在奖励,适合多智能体协作场景。

在人类社会中,个人利益与集体福祉之间的冲突常导致共同困境,如公地悲剧和社交困境。随着人工智能代理越来越多地作为人类的自主代理,我们提出一种新的多智能体强化学习(MARL)方法,旨在即使个体利益与集体目标冲突时,仍能最大化集体回报。不同于传统合作式MARL通过共享奖励、价值或策略,或设计内在奖励来促进协同学习,我们提出一种新方法:智能体之间交换行动建议。该方法相比共享奖励、价值或策略,泄露更少隐私信息,且无需设计内在奖励即可实现有效协作。我们的理论分析建立了集体与个体目标差异的上界,证明了建议共享能引导智能体行为趋同于集体目标。实验结果表明,该算法在性能上可媲美依赖值函数或策略共享、或使用内在奖励的基线方法。

原文摘要 · Abstract (English)

In human society, the conflict between self-interest and collective well-being often obstructs efforts to achieve shared welfare. Related concepts like the Tragedy of the Commons and Social Dilemmas frequently manifest in our daily lives. As artificial agents increasingly serve as autonomous proxies for humans, we propose a novel multi-agent reinforcement learning (MARL) method to address this issue - learning policies to maximise collective returns even when individual agents' interests conflict with the collective one. Unlike traditional cooperative MARL solutions that involve sharing rewards, values, and policies or designing intrinsic rewards to encourage agents to learn collectively optimal policies, we propose a novel MARL approach where agents exchange action suggestions. Our method reveals less private information compared to sharing rewards, values, or policies, while enabling effective cooperation without the need to design intrinsic rewards. Our algorithm is supported by our theoretical analysis that establishes a bound on the discrepancy between collective and individual objectives, demonstrating how sharing suggestions can align agents' behaviours with the collective objective. Experimental results demonstrate that our algorithm performs competitively with baselines that rely on value or policy sharing or intrinsic rewards.

多智能体强化学习协同决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。