SAIL-Embedding统一多模态表示,提升推荐系统长期体验与匹配效果。
SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model
- 分阶段训练+内容自适应策略,增强跨模态泛化能力。
- 在线实验显示抖音推荐场景7天留存提升0.5%,匹配准确率增0.1%。
- 专为工业级推荐场景设计,适合大规模多模态应用落地。
多模态嵌入模型旨在生成信息丰富的统一表示以支持多样化的跨模态任务。尽管从CLIP类双塔架构到大视觉语言模型的发展取得了进展,现有方法在真实应用场景和商业实践中仍面临模态支持有限、训练不稳定及领域差距等挑战。本文提出SAIL-Embedding,一种面向全模态的嵌入基础模型,通过定制化训练策略与架构设计解决上述问题。优化过程中,采用多阶段训练方案:内容感知渐进式训练提升模型对下游任务的适应性与跨模态能力;协作感知推荐增强训练通过融合序列到物品和ID到物品嵌入,挖掘用户历史兴趣,优化推荐表示。同时引入随机专业化与数据驱动模式匹配机制,提升训练灵活性与泛化性能。实验表明,SAIL-Embedding在各类检索任务中达到当前最优表现。在线实验显示,在多个真实场景集成后,用户生命周期(LT)显著提升,如在抖音精选场景中7天留存增长0.5%;对于抖音信息流排序模型,其生成的匹配特征带来0.1%的AUC提升。
原文摘要 · Abstract (English)
Multimodal embedding models aim to yield informative unified representations that empower diverse cross-modal tasks. Despite promising developments in the evolution from CLIP-based dual-tower architectures to large vision-language models, prior works still face unavoidable challenges in real-world applications and business scenarios, such as the limited modality support, unstable training mechanisms, and industrial domain gaps. In this work, we introduce SAIL-Embedding, an omni-modal embedding foundation model that addresses these issues through tailored training strategies and architectural design. In the optimization procedure, we propose a multi-stage training scheme to boost the multifaceted effectiveness of representation learning. Specifically, the content-aware progressive training aims to enhance the model's adaptability to diverse downstream tasks and master enriched cross-modal proficiency. The collaboration-aware recommendation enhancement training further adapts multimodal representations for recommendation scenarios by distilling knowledge from sequence-to-item and ID-to-item embeddings while mining user historical interests. Concurrently, we develop the stochastic specialization and dataset-driven pattern matching to strengthen model training flexibility and generalizability. Experimental results show that SAIL-Embedding achieves SOTA performance compared to other methods in different retrieval tasks. In online experiments across various real-world scenarios integrated with our model, we observe a significant increase in Lifetime (LT), which is a crucial indicator for the recommendation experience. For instance, the model delivers the 7-day LT gain of +0.5% in the Douyin-Selected scenario. For the Douyin feed rank model, the match features produced by SAIL-Embedding yield a +0.1% AUC gain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。