通过相对优势去偏,提升短视频推荐中的观看时长预测准确性
Relative Advantage Debiasing for Watch-Time Prediction in Short-Video Recommendation
- 用用户和视频分组的参考分布校正原始观看时长
- 两阶段架构分离分布估计与偏好学习,提升模型稳定性
- 引入分布嵌入,无需存储历史数据即可高效建模量化偏好
观看时长广泛用作视频推荐平台中用户满意度的代理指标。然而,原始观看时长受视频时长、流行度及个体用户行为等混淆因素影响,可能扭曲偏好信号,导致推荐模型偏差。本文提出一种新型相对优势去偏框架,通过对比用户与物品分组的实证参考分布来校正观看时长,生成基于分位数的偏好信号,并采用两阶段架构显式分离分布估计与偏好学习过程。此外,我们设计分布嵌入以高效参数化观看时长分位数,避免在线采样或历史数据存储。离线与在线实验均表明,该方法在推荐准确性和鲁棒性上显著优于现有基线方法。
原文摘要 · Abstract (English)
Watch time is widely used as a proxy for user satisfaction in video recommendation platforms. However, raw watch times are influenced by confounding factors such as video duration, popularity, and individual user behaviors, potentially distorting preference signals and resulting in biased recommendation models. We propose a novel relative advantage debiasing framework that corrects watch time by comparing it to empirically derived reference distributions conditioned on user and item groups. This approach yields a quantile-based preference signal and introduces a two-stage architecture that explicitly separates distribution estimation from preference learning. Additionally, we present distributional embeddings to efficiently parameterize watch-time quantiles without requiring online sampling or storage of historical data. Both offline and online experiments demonstrate significant improvements in recommendation accuracy and robustness compared to existing baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。