用统一框架同时提升广告排序模型的泛化能力和效率
A Unified Knowledge-Distillation and Semi-Supervised Learning Framework to Improve Industrial Ads Delivery Systems
- 融合知识蒸馏与半监督学习,利用海量无标签数据训练
- 在真实工业场景中实现多亿级用户、多场景下的性能提升
- 适合大规模广告系统优化,尤其关注数据偏差与模型泛化
工业广告排序系统传统上依赖标注的曝光数据,导致过拟合、模型扩展增益缓慢以及训练与线上数据差异带来的偏差。为此,我们提出统一的知识蒸馏与半监督学习框架(UKDSL),使模型能够基于更大更丰富的数据集进行训练,从而降低过拟合风险并缓解训练-服务数据不一致问题。我们提供了多阶段排序系统内在误校准与预测偏差的形式化分析及数值模拟,并通过实证证明该框架能有效缓解这些问题。相比已有方法,UKDSL可让模型从大量无标签数据中学习,显著提升性能且计算高效。最后,我们报告了UKDSL在工业环境中的成功部署,覆盖多亿级用户、多种服务界面、地理区域和客户,支持多种转化事件优化,据我们所知,这是首个在如此大规模与高效率下运行的同类框架。
原文摘要 · Abstract (English)
Industrial ads ranking systems conventionally rely on labeled impression data, which leads to challenges such as overfitting, slower incremental gain from model scaling, and biases due to discrepancies between training and serving data. To overcome these issues, we propose a Unified framework for Knowledge-Distillation and Semi-supervised Learning (UKDSL) for ads ranking, empowering the training of models on a significantly larger and more diverse datasets, thereby reducing overfitting and mitigating training-serving data discrepancies. We provide detailed formal analysis and numerical simulations on the inherent miscalibration and prediction bias of multi-stage ranking systems, and show empirical evidence of the proposed framework's capability to mitigate those. Compared to prior work, UKDSL can enable models to learn from a much larger set of unlabeled data, hence, improving the performance while being computationally efficient. Finally, we report the successful deployment of UKDSL in an industrial setting across various ranking models, serving users at multi-billion scale, across various surfaces, geological locations, clients, and optimize for various events, which to the best of our knowledge is the first of its kind in terms of the scale and efficiency at which it operates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。