arXiv:2410.18111cs.IRcs.LG2024-10被引 2

提升大模型推荐系统的数据效率,降低训练成本。

Data Efficiency for Large Recommendation Models

  • 提出数据收敛概念与加速方法
  • 实现训练数据量与模型规模的最优平衡
  • 适用于广告点击率预测等大规模场景

大规模推荐模型(LRMs)是数十亿美元在线广告产业的核心,需处理数百亿级别的数据样本,再通过持续在线训练适应快速变化的用户行为。数据规模直接影响计算成本和新方法的评估速度(研发速度)。本文提出可操作的原则与高层框架,指导从业者优化训练数据需求。这些策略已在谷歌最大广告点击率预测模型中成功部署,并广泛适用于其他大规模推荐系统。文章阐述了数据收敛的概念,描述了加速收敛的方法,并详细说明如何在训练数据量与模型规模之间实现最优权衡。

原文摘要 · Abstract (English)

Large recommendation models (LRMs) are fundamental to the multi-billion dollar online advertising industry, processing massive datasets of hundreds of billions of examples before transitioning to continuous online training to adapt to rapidly changing user behavior. The massive scale of data directly impacts both computational costs and the speed at which new methods can be evaluated (R&D velocity). This paper presents actionable principles and high-level frameworks to guide practitioners in optimizing training data requirements. These strategies have been successfully deployed in Google's largest Ads CTR prediction models and are broadly applicable beyond LRMs. We outline the concept of data convergence, describe methods to accelerate this convergence, and finally, detail how to optimally balance training data volume with model size.

推荐系统数据效率大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。