arXiv:2509.00199cs.IRcs.LG2025-09被引 1

在线实验中算法适应偏差会导致新模型被低估,影响推荐系统上线决策。

Algorithm Adaptation Bias in Recommendation System Online Experiments

  • 新模型在小流量测试中因用户数据受旧系统影响而表现失真。
  • 真实部署效果与实验结果可能严重偏离,大流量旧模型常被高估。
  • 提出需改进实验设计与评估方法,避免错过真正优秀的模型。

在线实验(A/B测试)是评估推荐系统变体和指导上线决策的金标准。然而,多种偏差可能扭曲实验结果并误导决策。一种未被充分研究但关键的偏差是算法适应效应。该偏差源于生产模型、用户数据与训练流程之间的飞轮动态:新模型在由现有系统塑造的数据分布上进行评估,或仅在小流量处理组中测试。结果导致新功能在受限实验环境中的建模与用户体验效果,与其全量部署的真实影响存在显著差异。实践中,实验结果往往偏好大流量的现有版本,低估小流量测试版本的表现,从而错失真正优胜方案或低估其影响。本文旨在引起对算法适应偏差的关注,将其置于推荐系统评估偏差的更广泛背景中,并推动关于跨实验设计、度量与调整解决方案的讨论。我们详细阐述该偏差机制,基于真实世界实验提供实证证据,并探讨更稳健的在线评估方法。

原文摘要 · Abstract (English)

Online experiments (A/B tests) are widely regarded as the gold standard for evaluating recommender system variants and guiding launch decisions. However, a variety of biases can distort the results of the experiment and mislead decision-making. An underexplored but critical bias is algorithm adaptation effect. This bias arises from the flywheel dynamics among production models, user data, and training pipelines: new models are evaluated on user data whose distributions are shaped by the incumbent system or tested only in a small treatment group. As a result, the measured effect of a new product change in modeling and user experience in this constrained experimental setting can diverge substantially from its true impact in full deployment. In practice, the experiment results often favor the production variant with large traffic while underestimating the performance of the test variant with small traffic, which leads to missing opportunities to launch a true winning arm or underestimating the impact. This paper aims to raise awareness of algorithm adaptation bias, situate it within the broader landscape of RecSys evaluation biases, and motivate discussion of solutions that span experiment design, measurement, and adjustment. We detail the mechanisms of this bias, present empirical evidence from real-world experiments, and discuss potential methods for a more robust online evaluation.

推荐系统在线实验偏差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。