arXiv:2502.07307cs.IR2025-02

用大模型模拟创作者行为,更真实评估推荐系统长期效果。

CreAgent: Towards Long-Term Evaluation of Recommender System under Platform-Creator Information Asymmetry

  • 基于大模型和博弈论,模拟创作者在信息不对称下的决策
  • 仿真结果与真实平台行为高度一致,提升评估可信度
  • 适合研究推荐系统公平性、多样性及长期可持续性的学者

推荐系统(RS)的长期可持续性成为关键问题。传统离线评估方法多关注用户即时反馈(如点击),却忽视内容创作者的长期影响。在真实内容平台中,创作者会根据用户反馈和趋势策略性地生产内容。现有研究虽尝试建模创作者行为,但普遍忽略平台与创作者间的信息不对称:创作者仅能获取自己作品的反馈,而平台掌握全局用户数据。当前推荐系统仿真器未考虑此不对称性,导致长期评估失真。为此,我们提出 CreAgent——一个由大语言模型驱动的创作者仿真代理。通过融合博弈论信念机制与快慢思维框架,CreAgent 能有效模拟信息不对称下的创作者行为。此外,我们使用近端策略优化(PPO)对模型进行微调,进一步提升其仿真能力。可信度验证实验表明,CreAgent 的行为与真实平台-创作者互动高度吻合,显著提升了长期评估的可靠性。通过引入 CreAgent 进行系统仿真,我们还可探究公平性与多样性感知算法如何促进各利益相关方的长期表现。CreAgent 及仿真平台已开源:https://github.com/shawnye2000/CreAgent。

原文摘要 · Abstract (English)

Ensuring the long-term sustainability of recommender systems (RS) emerges as a crucial issue. Traditional offline evaluation methods for RS typically focus on immediate user feedback, such as clicks, but they often neglect the long-term impact of content creators. On real-world content platforms, creators can strategically produce and upload new items based on user feedback and preference trends. While previous studies have attempted to model creator behavior, they often overlook the role of information asymmetry. This asymmetry arises because creators primarily have access to feedback on the items they produce, while platforms possess data on the entire spectrum of user feedback. Current RS simulators, however, fail to account for this asymmetry, leading to inaccurate long-term evaluations. To address this gap, we propose CreAgent, a Large Language Model (LLM)-empowered creator simulation agent. By incorporating game theory's belief mechanism and the fast-and-slow thinking framework, CreAgent effectively simulates creator behavior under conditions of information asymmetry. Additionally, we enhance CreAgent's simulation ability by fine-tuning it using Proximal Policy Optimization (PPO). Our credibility validation experiments show that CreAgent aligns well with the behaviors between real-world platform and creator, thus improving the reliability of long-term RS evaluations. Moreover, through the simulation of RS involving CreAgents, we can explore how fairness- and diversity-aware RS algorithms contribute to better long-term performance for various stakeholders. CreAgent and the simulation platform are publicly available at https://github.com/shawnye2000/CreAgent.

推荐系统创作者模拟信息不对称长期评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。