arXiv:2605.24516cs.MAcs.AI2026-05

动态调整惩罚力度,让合作更持久且不伤自身利益。

Adaptive Punishment for Cooperation in Mixed-Motive Games

论文配图:Adaptive Punishment for Cooperation in Mixed-Motive Games
图 1 · 摘自论文原文
  • 根据背叛严重程度与概率动态调节惩罚强度
  • 在迭代公共品博弈中显著提升合作率
  • 适合研究多智能体协作与社会困境的学者

现实中的多智能体交互普遍存在混合动机问题,自利个体为获取即时收益常选择背叛,忽视利他合作带来的长期收益与集体福祉。同伴惩罚可抑制背叛行为,但作为代价高昂的二次合作行为,持续施罚可能损害惩罚者自身利益。现有方法难以有效实施惩罚以促进合作。为此,我们提出自适应合作惩罚(APC),一种分布式方法,依据动态惩罚概率和背叛严重程度决定惩罚强度。该动态概率大幅减少无效且高成本的惩罚,同时促进合作。为准确评估背叛及其严重性,引入由游戏奖励引导学习的背叛感知模块。理论分析与实证结果表明,APC在迭代公共品博弈中表现优异;在多种顺序社会困境任务中,显著优于现有基线,学习到理性且高效的惩罚策略,通过战略性威慑背叛来促进合作。

原文摘要 · Abstract (English)

Mixed-motive scenarios are ubiquitous in real-world multi-agent interactions, where self-interested agents often defect for immediate rewards, overlooking the potential of altruistic cooperation to improve long-term gains and collective welfare. Peer punishment can deter defection, but as costly second-order altruism, its persistent imposition may undermine the punisher's interests. Existing approaches often struggle to effectively implement punishment to promote cooperation. To balance the efficacy and cost of punishment, we propose Adaptive Punishment for Cooperation (APC), a distributed method that determines punishment intensity based on both a dynamic punishment probability and the severity of defection. This dynamic probability substantially reduces costly and ineffective punishment while also promotes cooperation. To accurately assess defection and its severity, we use a defection awareness module, whose learning is guided by game reward. Theoretical analysis and empirical results show APC performs effectively in iterated public goods game. Empirically, APC also significantly outperforms existing baselines across sequential social dilemmas, learning rational and effective punishment policies that foster cooperation by strategically deterring defection.

多智能体合作机制社会困境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。