让智能体自愿承诺计划,提升合作效率
Learning to Negotiate via Voluntary Commitment
- 设计可学习的自愿承诺机制,促进智能体间合作
- 实验显示收敛更快,社会福利回报更高
- 适合多智能体协作场景,如资源分配与博弈
自主智能体间的部分对齐与冲突在诸多现实应用中形成混合动机场景。然而,即使合作能带来更好结果,智能体仍可能无法协同。一个关键原因是承诺不可信。为此,我们提出马尔可夫承诺游戏(MCG),允许智能体自愿承诺未来计划。基于此,我们设计了基于策略梯度的可学习承诺协议,并引入激励相容学习以加速收敛至高社会福利均衡。在具有挑战性的混合动机任务中,实验表明本方法相比基线实现更快的收敛速度和更高的收益。代码已公开于 https://github.com/shuhui-zhu/DCL。
原文摘要 · Abstract (English)
The partial alignment and conflict of autonomous agents lead to mixed-motive scenarios in many real-world applications. However, agents may fail to cooperate in practice even when cooperation yields a better outcome. One well known reason for this failure comes from non-credible commitments. To facilitate commitments among agents for better cooperation, we define Markov Commitment Games (MCGs), a variant of commitment games, where agents can voluntarily commit to their proposed future plans. Based on MCGs, we propose a learnable commitment protocol via policy gradients. We further propose incentive-compatible learning to accelerate convergence to equilibria with better social welfare. Experimental results in challenging mixed-motive tasks demonstrate faster empirical convergence and higher returns for our method compared with its counterparts. Our code is available at https://github.com/shuhui-zhu/DCL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。