用强化学习解释人类为何偏好公平,发现内在激励足以催生公平行为。
Decoding fairness: a reinforcement learning perspective
- 用双Q表的强化学习模拟议价双方,以累积奖励为目标决策。
- 高报价提升成交概率,与真实实验结果一致,公平策略自然涌现。
- 无需外因,内在机制即可稳定生成公平或理性策略,适用于多种场景。
对最后通牒博弈(UG)的行为实验表明,人类偏好公平行为,这与传统经济学预测相悖。现有解释多归因于模仿学习框架中的外生因素。本文采用强化学习范式,个体通过最大化累积奖励来决策。具体地,将Q-learning应用于UG,为每位玩家分配两个Q表,分别指导提议者和回应者的决策。在双人场景中,当同时考虑经验积累与未来回报时,公平性显著浮现。尤其当报价提高时,交易成功率随之上升,与行为实验观察一致。机制分析显示系统经历两个阶段,最终趋于公平或理性策略的稳定状态。该结论在角色轮换改为随机或固定分配,或扩展至格点群体时依然稳健。因此,研究结论表明,内生因素已足够解释公平性的出现,无需依赖外生因素。
原文摘要 · Abstract (English)
Behavioral experiments on the ultimatum game (UG) reveal that we humans prefer fair acts, which contradicts the prediction made in orthodox Economics. Existing explanations, however, are mostly attributed to exogenous factors within the imitation learning framework. Here, we adopt the reinforcement learning paradigm, where individuals make their moves aiming to maximize their accumulated rewards. Specifically, we apply Q-learning to UG, where each player is assigned two Q-tables to guide decisions for the roles of proposer and responder. In a two-player scenario, fairness emerges prominently when both experiences and future rewards are appreciated. In particular, the probability of successful deals increases with higher offers, which aligns with observations in behavioral experiments. Our mechanism analysis reveals that the system undergoes two phases, eventually stabilizing into fair or rational strategies. These results are robust when the rotating role assignment is replaced by a random or fixed manner, or the scenario is extended to a latticed population. Our findings thus conclude that the endogenous factor is sufficient to explain the emergence of fairness, exogenous factors are not needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。