用原型组合方式生成语义标识符,提升推荐系统点击率预测效果。
PaletteID: Prototype-Composed Semantic Identifiers for Multimodal CTR Prediction

- 基于原型池构建语义锚点,融合内容密度与语义多样性
- 在两个数据集上显著提升点击率预测性能,长尾物品收益更大
- 标识符更可解释,适合需要可解释推荐的场景
多模态信息能提升点击率(CTR)预测精度,有效缓解物品冷启动和长尾问题。现有方法通常将预训练多模态嵌入离散化为语义标识符(SIDs),使模型学习任务相关的语义表示。但现有方法仍受限于两大问题:一是码本分配未能保留语义相关性,丢弃原始嵌入空间中的细粒度连续信号;二是残差码路径高度依赖前缀码,限制了分层标识符的有效表征扩展性。为此,我们提出PaletteID(PID),一种基于原型的语义标识符。受调色板式色彩组合启发,PID使用一组代表性原型项作为语义锚点,连接预训练多模态内容空间与推荐模型。具体地,首先通过语义质量感知的行列式点过程(SQ-DPP)构建原型调色板,同时考虑局部内容密度与全局语义多样性;随后,对每个目标物品检索一系列语义相关的原型,并聚合形成具有信息量的PID表示,实现丰富且互补的语义建模。在两个公开数据集上的大量实验表明,PID持续提升CTR预测效果,对长尾物品的增益尤为显著。同时,PID生成的标识符分配更鲁棒,语义更可解释,优于现有残差式SID方法。
原文摘要 · Abstract (English)
Multimodal information can improve the accuracy of click-through rate (CTR) prediction and effectively alleviate item cold-start and long-tail problems. Recent studies commonly discretize pretrained multimodal embeddings into semantic identifiers (SIDs), allowing the model to learn task-specific semantic representations for recommendation. However, existing methods still provide limited gains due to two major limitations. First, codebook assignment fails to preserve semantic relevance and discards fine-grained continuous signals in the original embedding space. Second, the residual code paths are highly dependent on prefix codes, which limits the effective representational scalability of hierarchical identifiers. To address these issues, we propose PaletteID (PID), a prototype-based semantic identifier. Inspired by palette-based color composition, PID uses a compact set of representative prototype items as semantic anchors to bridge pretrained multimodal content space and recommendation models. Specifically, we first construct a prototype palette with Semantic Quality-Aware Determinantal Point Process (SQ-DPP), which jointly considers local content density and global semantic diversity. Then, for each target item, PID retrieves a sequence of semantically related prototypes and aggregates them into an informative PID representation, enabling rich and complementary semantic modeling. Extensive experiments on two public datasets demonstrate that PID consistently improves CTR prediction and yields larger gains for long-tail items. PID also produces more robust identifier assignments and provides more interpretable token semantics than existing residual SID methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。