用不确定性建模修复短视频推荐中的满意度标签偏差
Uncertainty as Remedy: Mitigating Satisfaction Label Bias in Short Video Multi-Objective Ensemble Ranking

- 将预测结果建模为高斯分布,用方差表示预测不确定性
- 通过概率成对损失和样本加权,有效降低标签偏差
- 在真实工业平台验证,显著提升推荐效果与用户满意度一致性
短视频推荐的核心目标是建模用户对推荐视频的隐含真实满意度。当前主流的端到端多目标集成排序模型通常基于点击、观看时长等多维密集行为信号训练,但这些信号具有局部性、碎片化且常相互冲突,引入了不确定性与标签偏差。传统确定性模型忽略此问题,加剧了偏差并导致模型收敛不佳。现有不确定性方法多用于事后排序调整,而非内嵌于核心优化流程中。本文提出UAME框架,将模型预测表示为高斯评分变量,均值为预测满意度,方差表征对应不确定性。设计概率成对排序损失,并构建样本级不确定性加权机制以缓解偏差。理论分析表明该加权策略有助于减轻满意度标签偏差。大规模离线与在线实验显示,UAME持续改进两种先进范式EMER与EASQ,更贴近问卷调查的用户满意度。已部署于生产系统,持续带来稳定显著收益。
原文摘要 · Abstract (English)
The core objective of short video recommendation is to model users' unobservable true satisfaction with recommended videos. As the dominant industrial framework, end-to-end multi-objective ensemble ranking models are typically trained with multi-dimensional dense user behavioral signals, such as clicks and watch time. However, these behavioral signals are partial, fragmented, and often mutually conflicting user satisfaction proxies, introducing uncertainty and label bias into satisfaction modeling. Conventional deterministic models overlook this uncertainty, which exacerbates satisfaction label bias and results in suboptimal model convergence. Meanwhile, existing uncertainty-aware methods mostly employ uncertainty for post-hoc ranking adjustments rather than leveraging it as a remedy to mitigate the inherent bias within the core optimization pipeline. This paper proposes UAME, an Uncertainty-Aware end-to-end Multi-objective Ensemble ranking framework for short video recommendation. UAME represents the model's prediction as a Gaussian scoring variable, where the mean denotes the predicted satisfaction score and the variance quantifies predictive uncertainty associated with this score. We further design a probabilistic pairwise ranking loss, and construct an uncertainty-aware sample-level weighting scheme to mitigate the bias. We further provide theoretical analysis suggesting that the weighting scheme helps mitigate satisfaction label bias. Extensive offline and online experiments on a large-scale industrial short video platform demonstrate that UAME consistently improves two state-of-the-art paradigms, EMER and EASQ, and better aligns with questionnaire-based user satisfaction. UAME has been deployed in our production short-video recommendation system and continues to deliver stable, statistically significant gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。