用大模型模拟真实用户行为,提升推荐系统评估准确性
SimUSER: Simulating User Behavior with Large Language Models for Recommender System Evaluation
- 构建带人格、记忆和感知的虚拟用户代理
- 模拟结果与真实用户行为高度吻合,点击率预测更准
- 适合推荐系统研发与评测团队使用
推荐系统在众多实际应用中发挥核心作用,但其性能评估仍面临离线指标与在线行为之间的差距。由于真实用户数据稀缺且存在隐私限制,我们提出 SimUSER,一个可信赖且低成本的人类代理框架。该框架首先从历史数据中识别出自洽的用户人格,丰富用户画像的背景与个性特征。随后,关键在于配备人格、记忆、感知与决策模块的虚拟用户与推荐系统进行交互。SimUSER 在微观与宏观层面均表现出比以往工作更贴近真实人类的行为。我们还开展深入实验,探究缩略图对点击率的影响、曝光效应以及评论对用户参与度的作用。最终,基于离线 A/B 测试结果优化推荐系统参数,显著提升了真实场景中的用户参与度。
原文摘要 · Abstract (English)
Recommender systems play a central role in numerous real-life applications, yet evaluating their performance remains a significant challenge due to the gap between offline metrics and online behaviors. Given the scarcity and limits (e.g., privacy issues) of real user data, we introduce SimUSER, an agent framework that serves as believable and cost-effective human proxies. SimUSER first identifies self-consistent personas from historical data, enriching user profiles with unique backgrounds and personalities. Then, central to this evaluation are users equipped with persona, memory, perception, and brain modules, engaging in interactions with the recommender system. SimUSER exhibits closer alignment with genuine humans than prior work, both at micro and macro levels. Additionally, we conduct insightful experiments to explore the effects of thumbnails on click rates, the exposure effect, and the impact of reviews on user engagement. Finally, we refine recommender system parameters based on offline A/B test results, resulting in improved user engagement in the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。