用仿真模型替代部分真实实验,高效评估新用户推荐策略效果
Minimizing Live Experiments in Recommender Systems: User Simulation to Evaluate Preference Elicitation Policies
- 构建反事实鲁棒的用户行为模型,模拟真实用户决策过程
- 仿真结果与上线后关键指标高度一致,预测准确率提升显著
- 适合需要快速验证推荐策略的平台研发团队使用
推荐系统策略的评估通常依赖于对真实用户的实时A/B测试,但这种方法周期长、成本高,且可能影响用户留存。尤其在新用户引导阶段,由于仅发生一次,成本问题更为突出。本文提出一种仿真方法,用于补充并减少对真实实验的依赖。通过构建反事实鲁棒的用户行为模型,并与生产系统集成,实现了对新用户偏好获取算法的有效评估。该方法在YouTube Music平台上验证,能可靠预测算法上线后的关键指标表现。文中详述了应用场景、仿真模型与平台架构、实验结果及部署成效,并指出未来需进一步提升仿真真实性以增强其作为真实实验有力补充的能力。
原文摘要 · Abstract (English)
Evaluation of policies in recommender systems typically involves A/B testing using live experiments on real users to assess a new policy's impact on relevant metrics. This ``gold standard'' comes at a high cost, however, in terms of cycle time, user cost, and potential user retention. In developing policies for ``onboarding'' new users, these costs can be especially problematic, since on-boarding occurs only once. In this work, we describe a simulation methodology used to augment (and reduce) the use of live experiments. We illustrate its deployment for the evaluation of ``preference elicitation'' algorithms used to onboard new users of the YouTube Music platform. By developing counterfactually robust user behavior models, and a simulation service that couples such models with production infrastructure, we are able to test new algorithms in a way that reliably predicts their performance on key metrics when deployed live. We describe our domain, our simulation models and platform, results of experiments and deployment, and suggest future steps needed to further realistic simulation as a powerful complement to live experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。