用统一语义ID提升生成式搜索与推荐的协同效果
Semantic IDs for Joint Generative Search and Recommendation
- 用跨任务微调的双编码器生成统一语义ID
- 联合优化后在搜索和推荐上均表现优异
- 适合构建下一代统一生成式推荐系统
由大语言模型驱动的生成式模型正成为统一支持推荐与搜索任务的方案。其关键设计在于如何表示物品:传统使用唯一标识符(ID),近年则采用由嵌入向量生成的离散代码形式的语义ID。尽管特定任务的嵌入模型可提升单一任务性能,但在联合设置下泛化能力有限。本文探讨如何构建在联合搜索与推荐中表现良好的语义ID。我们比较了多种构建策略,包括任务特定与跨任务方法,以及是否为每个任务单独分配语义ID令牌。结果表明,对搜索与推荐任务联合微调的双编码器模型生成的嵌入,再构建统一的语义ID空间,能实现良好平衡,显著提升两项任务的表现。这些发现有望推动可泛化的语义化ID体系发展,并为下一代统一生成式推荐架构提供参考。
原文摘要 · Abstract (English)
Generative models powered by Large Language Models (LLMs) are emerging as a unified solution for powering both recommendation and search tasks. A key design choice in these models is how to represent items, traditionally through unique identifiers (IDs) and more recently with Semantic IDs composed of discrete codes, obtained from embeddings. While task-specific embedding models can improve performance for individual tasks, they may not generalize well in a joint setting. In this paper, we explore how to construct Semantic IDs that perform well both in search and recommendation when using a unified model. We compare a range of strategies to construct Semantic IDs, looking into task-specific and cross-tasks approaches, and also whether each task should have its own semantic ID tokens in a joint search and recommendation generative model. Our results show that using a bi-encoder model fine-tuned on both search and recommendation tasks to obtain item embeddings, followed by the construction of a unified Semantic ID space provides an effective trade-off, enabling strong performance in both tasks. We hope these findings spark follow-up work on generalisable, semantically grounded ID schemes and inform the next wave of unified generative recommender architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。