通过系统级协同设计,实现推荐模型训练快5倍、推理快21倍。
Bending the Scaling Law Curve in Large-Scale Recommendation Systems
- 创新输入序列与稀疏注意力机制,突破传统自注意力瓶颈。
- 训练速度提升5倍以上,推理速度提升21倍,效果更优。
- 已落地亿级用户场景,带来4%-8%的使用时长与互动增长。
基于用户交互历史的序列建模已成为大规模推荐系统的核心。大语言模型的进展揭示了有前景的扩展规律,推动了长序列建模与更深架构在推荐任务中的研究。然而,许多近期方法依赖交叉注意力缓解序列建模中的二次计算瓶颈,可能限制自注意力带来的表征能力。我们提出ULTRA-HSTU,一种通过端到端模型与系统协同设计的新型序列推荐模型。通过输入序列设计、稀疏注意力机制与模型拓扑的创新,该模型在模型质量与效率上均取得显著提升。全面基准测试表明,相比传统模型,ULTRA-HSTU实现超过5倍的训练加速与21倍的推理加速,同时提供更优的推荐效果。该方案已在实际生产中规模化部署,每日服务数十亿用户,在真实环境中带来4%至8%的消费与互动提升。
原文摘要 · Abstract (English)
Learning from user interaction history through sequential models has become a cornerstone of large-scale recommender systems. Recent advances in large language models have revealed promising scaling laws, sparking a surge of research into long-sequence modeling and deeper architectures for recommendation tasks. However, many recent approaches rely heavily on cross-attention mechanisms to address the quadratic computational bottleneck in sequential modeling, which can limit the representational power gained from self-attention. We present ULTRA-HSTU, a novel sequential recommendation model developed through end-to-end model and system co-design. By innovating in the design of input sequences, sparse attention mechanisms, and model topology, ULTRA-HSTU achieves substantial improvements in both model quality and efficiency. Comprehensive benchmarking demonstrates that ULTRA-HSTU achieves remarkable scaling efficiency gains -- over 5x faster training scaling and 21x faster inference scaling compared to conventional models -- while delivering superior recommendation quality. Our solution is fully deployed at scale, serving billions of users daily and driving significant 4% to 8% consumption and engagement improvements in real-world production environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。