arXiv:2605.21752cs.LGcs.AI2026-05

用对比学习纠正直播推荐中的用户活跃度偏差,提升推荐公平性与效果。

PEARL: Unbiased Percentile Estimation via Contrastive Learning for Industrial-Scale Livestream Recommendation

  • 通过对比真实交互样本直接估算相对偏好百分位,避免依赖分布模型。
  • 线上测试显示观看时长提升2.10%,互动率增1.49%,举报率降6.91%。
  • 适合大规模工业级推荐系统,尤其对高活跃用户主导的场景有显著改进。

基于用户行为数据训练的推荐系统易受行为强度不平衡影响——不同用户参与度差异导致反馈信号失真,使模型过度放大高活跃用户信号而忽略其他人,最终降低推荐质量与鲁棒性。为此,我们提出非参数对比百分位近似框架PEARL,通过建模相对偏好信号而非绝对活跃度来缓解此问题。基于相对优势去偏思想,PEARL利用真实对比交互样本来直接逼近百分位关系,无需辅助分布估计模型。理论上证明,成对比较可获得无偏的百分位偏好估计。为增强适用性,引入基于预测的自举平滑机制以处理稀疏离散反馈,并设计广义加权公式与联合训练策略,提升建模灵活性与表征学习能力。大量离线实验表明,PEARL有效缓解行为偏差,持续提升多目标排序性能。在日均用户超十亿的生产级直播平台部署后,线上A/B测试验证显著收益:观看时长+2.10%,消费金额+0.80%,互动率+1.49%,举报率-6.91%。

原文摘要 · Abstract (English)

Recommender systems trained on user interaction data are susceptible to behavioral intensity imbalance--a systematic distortion arising from heterogeneous engagement patterns across users. This imbalance skews feedback signals such that observed interactions no longer faithfully reflect true preferences, causing models to disproportionately amplify signals from highly active users while underrepresenting others, which ultimately degrades recommendation quality and robustness at scale. To address this issue, we propose a nonparametric contrastive percentile approximation framework, PEARL, that models relative preference signals instead of absolute engagement magnitudes. Building upon relative advantage debiasing, PEARL leverages real contrastive interaction samples to approximate percentile relationships directly, without relying on auxiliary distribution estimation models. We provide theoretical justification demonstrating that such pairwise comparisons yield unbiased estimates of percentile-based preference signals. For broader applicability, we introduce a prediction-based bootstrapping mechanism for percentile smoothing to handle sparse and discrete feedback, alongside a generalized value-weighted formulation and a co-training strategy to enhance both modeling flexibility and representation learning. Extensive offline experiments demonstrate that PEARL effectively mitigates behavioral bias and consistently improves recommendation performance across multiple ranking targets. Deployed in a production livestream platform with a combined user base of billions, online A/B testing confirms substantial real-world gains: +2.10% Watch Duration, +0.80% Consumption Amount, +1.49% Interaction Rate, and -6.91% Report Rate.

推荐系统去偏学习对比学习直播推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。