提升双塔推荐模型效率与效果,支持在线学习
CS3: Efficient Online Capability Synergy for Two-Tower Recommendation

- 通过自修正、塔间同步和级联共享三机制增强双塔模型
- 在三个公开数据集上优于强基线,广告系统中提升最高8.36%收益
- 轻量级设计适配线上学习,毫秒级延迟保持实时性
为平衡推荐系统的有效性和效率,多阶段流水线常采用轻量级双塔模型进行大规模候选召回。然而,孤立的双塔架构限制了表征能力、嵌入空间对齐和跨特征交互。现有方法如后期交互和知识蒸馏虽可缓解问题,但常增加延迟或难以部署于在线学习场景。我们提出高效在线框架能力协同(CS3),在保持实时约束的同时强化双塔检索器。CS3引入三项机制:(1) 周期自适应结构,通过各塔内自适应特征去噪实现自我修正;(2) 双塔同步,通过轻量级相互感知提升对齐;(3) 级联模型共享,通过复用下游模型知识增强跨阶段一致性。CS3可即插即用,兼容多种双塔骨干网络与在线学习。在三个公开数据集上的实验显示持续优于强基线,部署于大规模广告系统后,在三个场景中实现最高8.36%收入提升,同时维持毫秒级延迟。
原文摘要 · Abstract (English)
To balance effectiveness and efficiency in recommender systems, multi-stage pipelines commonly use lightweight two-tower models for large-scale candidate retrieval. However, the isolated two-tower architecture restricts representation capacity, embedding-space alignment, and cross-feature interactions. Existing solutions such as late interaction and knowledge distillation can mitigate these issues, but often increase latency or are difficult to deploy in online learning settings. We propose Capability Synergy (CS3), an efficient online framework that strengthens two-tower retrievers while preserving real-time constraints. CS3 introduces three mechanisms: (1) Cycle-Adaptive Structure for self-revision via adaptive feature denoising within each tower; (2) Cross-Tower Synchronization to improve alignment through lightweight mutual awareness between towers; and (3) Cascade-Model Sharing to enhance cross-stage consistency by reusing knowledge from downstream models. CS3 is plug-and-play with diverse two-tower backbones and compatible with online learning. Experiments on three public datasets show consistent gains over strong baselines, and deployment in a largescale advertising system yields up to 8.36% revenue improvement across three scenarios while maintaining ms-level latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。