arXiv:2504.03997cs.IR2025-04被引 1

用因果与信息论方法,让离线评估更公平准确。

Towards Robust Offline Evaluation: A Causal and Information Theoretic Framework for Debiasing Ranking Systems

  • 通过重加权将选择性数据转为随机缺失,减少偏差。
  • 在多个真实数据集上显著提升评估准确性,避免热门项误导。
  • 无需系统内部信息,适合各类推荐系统快速部署。

检索排序系统的评估对模型优化至关重要。尽管在线A/B测试是金标准,但其高成本和对用户体验的风险,促使我们发展有效的离线评估方法。然而,依赖历史交互数据会引入选择、曝光、从众及位置偏差,这些偏差源于用户行为的非随机缺失(MNAR)特性,导致热门或频繁曝光的项目被过度青睐,掩盖真实偏好。本文提出一种新的鲁棒离线评估框架,通过重加权结合黑盒优化,将MNAR数据转化为缺失随机(MAR),并以神经网络估计信息论指标进行引导。主要贡献包括:(1) 提出用于解决离线评估偏差的因果建模方法;(2) 构建系统无关的去偏框架;(3) 实证验证其有效性。该框架使评估更准确、公平且可泛化,显著提升模型部署前的评估质量。

原文摘要 · Abstract (English)

Evaluating retrieval-ranking systems is crucial for developing high-performing models. While online A/B testing is the gold standard, its high cost and risks to user experience require effective offline methods. However, relying on historical interaction data introduces biases-such as selection, exposure, conformity, and position biases-that distort evaluation metrics, driven by the Missing-Not-At-Random (MNAR) nature of user interactions and favoring popular or frequently exposed items over true user preferences. We propose a novel framework for robust offline evaluation of retrieval-ranking systems, transforming MNAR data into Missing-At-Random (MAR) through reweighting combined with black-box optimization, guided by neural estimation of information-theoretic metrics. Our contributions include (1) a causal formulation for addressing offline evaluation biases, (2) a system-agnostic debiasing framework, and (3) empirical validation of its effectiveness. This framework enables more accurate, fair, and generalizable evaluations, enhancing model assessment before deployment.

离线评估去偏因果推断推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。