arXiv:2510.10511cs.IR2025-10

通过信息披露引导创作者行为,提升推荐系统的长期用户福祉。

Enhancing Long-Term Welfare in Recommender Systems: An Information Revelation Approach

  • 用贝叶斯劝说构建平台与创作者间的信源-接收关系,实现行为引导。
  • 在两个真实数据集上显著优于现有公平重排序和信息披露方法。
  • 适合关注平台生态健康、长期用户留存的研究者与工程师。

提升推荐系统的长期用户福祉(如持续用户参与度)已成为核心目标。在现实平台中,内容创作者的生产行为对长期福祉的影响远超短期推荐精度,因此有效引导创作者行为对构建更健康的推荐生态至关重要。现有方法多依赖启发式重排序算法调整物品曝光以影响创作者,但此类策略常与短期推荐准确率目标冲突,导致性能下降,难以实现最优长期福祉。经济学研究提示:向信息较少的一方(接收者)披露信息丰富的另一方(发送者)的信息,可有效改变接收者的信念并引导其行为。受此启发,本文提出基于信息披露的长期福祉优化框架(LoRe)。将平台视为信息发送者,创作者为接收者,采用经典贝叶斯劝说方法建模。针对传统经济方法中不切实际的假设,将信息披露过程建模为马尔可夫决策过程(MDP),并设计在有限理性创作者环境中的学习与推理算法。在两个真实推荐系统数据集上的大量实验表明,该方法能显著优于现有公平重排序与信息披露策略,有效提升长期用户福祉。

原文摘要 · Abstract (English)

Improving the long-term user welfare (e.g., sustained user engagement) has become a central objective of recommender systems (RS). In real-world platforms, the creation behaviors of content creators plays a crucial role in shaping long-term welfare beyond short-term recommendation accuracy, making the effective steering of creator behavior essential to foster a healthier RS ecosystem. Existing works typically rely on re-ranking algorithms that heuristically adjust item exposure to steer creators' behavior. However, when embedded within recommendation pipelines, such a strategy often conflicts with the short-term objective of improving recommendation accuracy, leading to performance degradation and suboptimal long-term welfare. The well-established economics studies offer us valuable insights for an alternative approach without relying on recommendation algorithmic design: revealing information from an information-rich party (sender) to a less-informed party (receiver) can effectively change the receiver's beliefs and steer their behavior. Inspired by this idea, we propose an information-revealing framework, named Long-term Welfare Optimization via Information Revelation (LoRe). In this framework, we utilize a classical information revelation method (i.e., Bayesian persuasion) to map the stakeholders in RS, treating the platform as the sender and creators as the receivers. To address the challenge posed by the unrealistic assumption of traditional economic methods, we formulate the process of information revelation as a Markov Decision Process (MDP) and propose a learning algorithm trained and inferred in environments with boundedly rational creators. Extensive experiments on two real-world RS datasets demonstrate that our method can effectively outperform existing fair re-ranking methods and information revealing strategies in improving long-term user welfare.

推荐系统长期福祉信息披露贝叶斯劝说

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。