提出统一框架,让广告推荐的语义编号生成更准确。
End-to-End Semantic ID Generation for Generative Advertisement Recommendation
- 端到端联合优化嵌入与语义编号,避免两阶段压缩失真。
- 多粒度对比学习提升细粒度语义对齐,关键指标提升4.62%。
- 适合做广告推荐系统优化的研究者和工程师参考。
生成式推荐(GR)通过将推荐问题建模为下一步词预测,在广告推荐中表现优异。该范式依赖语义编号(SID)将大规模物品离散化为序列。现有方法主要使用残差量化(RQ)生成SID,即先编码物品为嵌入,再量化为离散编号。但该方法存在固有缺陷:1)两阶段压缩导致目标不一致与语义退化;2)RQ结构引发误差累积。为此,我们提出UniSID,一种面向生成式广告推荐的统一语义编号生成框架。具体而言,从原始广告数据端到端联合优化嵌入与SID,使语义信息直接流入编号空间,解决两阶段级联压缩的内在局限。为捕捉细粒度语义,引入多粒度对比学习策略,实现不同层级编号间的物品对齐。最后,设计基于摘要的广告重建机制,促使SID捕获广告上下文中未显式表达的高层语义。实验表明,UniSID在下游广告场景中持续优于现有最先进方法,命中率(Hit Rate)指标相较最强基线最高提升4.62%。
原文摘要 · Abstract (English)
Generative Recommendation (GR) has excelled by framing recommendation as next-token prediction. This paradigm relies on Semantic IDs (SIDs) to tokenize large-scale items into discrete sequences. Existing GR approaches predominantly generate SIDs via Residual Quantization (RQ), where items are encoded into embeddings and then quantized to discrete SIDs. However, this paradigm suffers from inherent limitations: 1) Objective misalignment and semantic degradation stemming from the two-stage compression; 2) Error accumulation inherent in the structure of RQ. To address these limitations, we propose UniSID, a Unified SID generation framework for generative advertisement recommendation. Specifically, we jointly optimize embeddings and SIDs in an end-to-end manner from raw advertising data, enabling semantic information to flow directly into the SID space and thus addressing the inherent limitations of the two-stage cascading compression paradigm. To capture fine-grained semantics, a multi-granularity contrastive learning strategy is introduced to align distinct items across SID levels. Finally, a summary-based ad reconstruction mechanism is proposed to encourage SIDs to capture high-level semantic information that is not explicitly present in advertising contexts. Experiments demonstrate that UniSID consistently outperforms state-of-the-art SID generation methods, yielding up to a 4.62% improvement in Hit Rate metrics across downstream advertising scenarios compared to the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。