arXiv:2508.05398cs.IRcs.LG2025-08中稿 · RecSys 2025被引 2

研究离线推荐评估中采样策略的可靠性,揭示其如何影响模型比较结果。

On the Reliability of Sampling Strategies in Offline Recommender Evaluation

  • 通过模拟不同曝光偏差,系统测试多种采样策略的性能。
  • 发现采样策略在模型区分度和结果稳定性上差异显著,影响评估可信度。
  • 为选择可靠、稳健的评估方法提供实证指导,适合评估者与研究者参考。

离线评估在在线测试不可行或存在风险时,是评测推荐系统的核心手段。然而,它易受两种主要偏差影响:暴露偏差(用户仅与被展示的项目互动)和采样偏差(评估基于日志中的子集而非完整商品目录)。尽管先前工作提出了缓解采样偏差的方法,但这些方法通常在固定日志数据集上评估,而非考察其在不同暴露条件下的模型比较可靠性或与真实用户偏好的一致程度。本文利用一个全量观测数据集作为真实基准,系统地模拟多种暴露偏差,并从四个维度评估常见采样策略的可靠性:采样分辨率(模型可区分性)、保真度(与完整评估的一致性)、鲁棒性(对暴露偏差的稳定性)以及预测能力(与真实基准的匹配度)。研究结果揭示了采样如何扭曲评估结果,并提供了选择能产生忠实且稳健比较结果的策略的实际建议。

原文摘要 · Abstract (English)

Offline evaluation plays a central role in benchmarking recommender systems when online testing is impractical or risky. However, it is susceptible to two key sources of bias: exposure bias, where users only interact with items they are shown, and sampling bias, introduced when evaluation is performed on a subset of logged items rather than the full catalog. While prior work has proposed methods to mitigate sampling bias, these are typically assessed on fixed logged datasets rather than for their ability to support reliable model comparisons under varying exposure conditions or relative to true user preferences. In this paper, we investigate how different combinations of logging and sampling choices affect the reliability of offline evaluation. Using a fully observed dataset as ground truth, we systematically simulate diverse exposure biases and assess the reliability of common sampling strategies along four dimensions: sampling resolution (recommender model separability), fidelity (agreement with full evaluation), robustness (stability under exposure bias), and predictive power (alignment with ground truth). Our findings highlight when and how sampling distorts evaluation outcomes and offer practical guidance for selecting strategies that yield faithful and robust offline comparisons.

推荐系统离线评估采样偏差实验设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。