用概率模型提升多模态嵌入离散化,显著改善点击率预测
RQ-GMM: Residual Quantized Gaussian Mixture Model for Multimodal Semantic Discretization in CTR Prediction
- 基于高斯混合与残差量化,建模嵌入空间统计结构
- 在公开数据集上比基线提升1.502%广告价值
- 适合大规模推荐系统中多模态特征处理
多模态内容对点击率(CTR)预测至关重要。然而,直接将预训练模型的连续嵌入输入到CTR模型中会因优化目标不一致和联合训练收敛速度差异导致效果不佳。在输入CTR模型前将嵌入离散化为语义ID是一种更有效的方法,但现有方法存在码本利用率低、重建精度不足和语义区分度弱的问题。我们提出RQ-GMM(残差量化高斯混合模型),引入概率建模以更好捕捉多模态嵌入空间的统计结构。通过高斯混合模型结合残差量化,RQ-GMM实现了更高的码本利用率和重建精度。在公开数据集上的实验及在服务数亿用户的大型短视频平台的在线A/B测试表明,RQ-GMM相比强基线带来1.502%的广告价值提升。该方法已全面部署,支撑每日推荐系统。
原文摘要 · Abstract (English)
Multimodal content is crucial for click-through rate (CTR) prediction. However, directly incorporating continuous embeddings from pre-trained models into CTR models yields suboptimal results due to misaligned optimization objectives and convergence speed inconsistency during joint training. Discretizing embeddings into semantic IDs before feeding them into CTR models offers a more effective solution, yet existing methods suffer from limited codebook utilization, reconstruction accuracy, and semantic discriminability. We propose RQ-GMM (Residual Quantized Gaussian Mixture Model), which introduces probabilistic modeling to better capture the statistical structure of multimodal embedding spaces. Through Gaussian Mixture Models combined with residual quantization, RQ-GMM achieves superior codebook utilization and reconstruction accuracy. Experiments on public datasets and online A/B tests on a large-scale short-video platform serving hundreds of millions of users demonstrate substantial improvements: RQ-GMM yields a 1.502% gain in Advertiser Value over strong baselines. The method has been fully deployed, serving daily recommendations for hundreds of millions of users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。