用博弈论设计公平激励机制,让内容创作者合作更稳定。
Creator Incentives in Recommender Systems: A Cooperative Game-Theoretic Approach for Stable and Fair Collaboration in Multi-Agent Bandits
- 将创作者合作建模为可转移效用的多人强化学习博弈
- 同质创作者下游戏凸性保证公平且稳定的激励分配
- 提出基于后悔值的分钱规则,适合实际推荐系统部署
在线推荐平台中,用户对某一创作者内容的反馈会影响系统的整体学习,进而影响其他创作者的内容曝光。为分析此类情境下的激励问题,本文将协作建模为带有可转移效用(TU)的多智能体随机线性伯努利问题,其中联盟的价值等于其成员累计后悔值的负和。在满足温和算法条件下,对于具有固定动作集的同质(相同)智能体,诱导出的TU博弈是凸的,意味着核心非空,包含谢尔普利值,从而保证了稳定性与公平性。对于异质智能体,尽管核心仍非空,但凸性及谢尔普利值属于核心不再成立。为此,本文提出一种基于后悔值的简单分配规则,满足谢尔普利公理中的三条,并位于核心内。在MovieLens-100k数据集上的实验展示了不同设置和算法下,实证分配结果与谢尔普利公平性的吻合与偏离情况。
原文摘要 · Abstract (English)
User interactions in online recommendation platforms create interdependencies among content creators: feedback on one creator's content influences the system's learning and, in turn, the exposure of other creators' contents. To analyze incentives in such settings, we model collaboration as a multi-agent stochastic linear bandit problem with a transferable utility (TU) cooperative game formulation, where a coalition's value equals the negative sum of its members' cumulative regrets. We show that, for identical (homogenous) agents with fixed action sets, the induced TU game is convex under mild algorithmic conditions, implying a non-empty core that contains the Shapley value and ensures both stability and fairness. For heterogeneous agents, the game still admits a non-empty core, though convexity and Shapley value core-membership are no longer guaranteed. To address this, we propose a simple regret-based payout rule that satisfies three out of the four Shapley axioms and also lies in the core. Experiments on MovieLens-100k dataset illustrate when the empirical payout aligns with -- and diverges from -- the Shapley fairness across different settings and algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。