用公平性替代功利主义,让多智能体合作更公正
Altruism and Fair Objective in Mixed-Motive Markov games
- 用比例公平代替功利主义目标,设计公平利他效用函数
- 在经典社会困境中推导出确保合作的解析条件
- 提出公平马尔可夫游戏与新算法,适合追求公平协作的研究者
合作是社会存续的基础,能使异质群体实现集体福祉。然而个体倾向于不付出成本而享受合作收益,导致不公平局面。博弈论中的社会困境正体现了个人利益与集体结果之间的矛盾。当前主流多智能体合作方法采用功利主义福利,虽高效但极度不公。本文提出新框架,以比例公平替代标准功利目标,为每个智能体定义基于个体对数收益空间的公平利他效用,并推导出在经典社会困境中确保合作的解析条件。进一步将该框架扩展至序列场景,定义公平马尔可夫游戏,并提出新型公平演员-评论家算法来学习公平策略。最后在多种社会困境环境上评估了该方法。
原文摘要 · Abstract (English)
Cooperation is fundamental for society's viability, as it enables the emergence of structure within heterogeneous groups that seek collective well-being. However, individuals are inclined to defect in order to benefit from the group's cooperation without contributing the associated costs, thus leading to unfair situations. In game theory, social dilemmas entail this dichotomy between individual interest and collective outcome. The most dominant approach to multi-agent cooperation is the utilitarian welfare which can produce efficient highly inequitable outcomes. This paper proposes a novel framework to foster fairer cooperation by replacing the standard utilitarian objective with Proportional Fairness. We introduce a fair altruistic utility for each agent, defined on the individual log-payoff space and derive the analytical conditions required to ensure cooperation in classic social dilemmas. We then extend this framework to sequential settings by defining a Fair Markov Game and deriving novel fair Actor-Critic algorithms to learn fair policies. Finally, we evaluate our method in various social dilemma environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。