arXiv:2410.09560cs.IRcs.LG2024-10被引 5

用多码本提升推荐系统的语义表示能力

Towards Scalable Semantic Representation for Recommendation

  • 构建多个独立码本压缩大模型语义嵌入
  • 在推荐任务中实现更强的区分度与稳定性
  • 适合需要高维语义表示的推荐系统研究者

随着大语言模型的发展,基于LLM生成语义ID以增强推荐系统性能的研究日益增多。然而,这些嵌入的维度需与推荐系统中ID嵌入的维度匹配,通常远小于原始长度,导致语义嵌入的区分度和维度鲁棒性不可避免地损失。为此,本文提出Mixture-of-Codes:在索引阶段构建多个独立码本对LLM表示进行编码,并在下游推荐阶段结合语义表示与融合模块。大量分析与实验表明,该方法显著提升了区分度与维度鲁棒性,实现了推荐场景下的最优可扩展性表现。

原文摘要 · Abstract (English)

With recent advances in large language models (LLMs), there has been emerging numbers of research in developing Semantic IDs based on LLMs to enhance the performance of recommendation systems. However, the dimension of these embeddings needs to match that of the ID embedding in recommendation, which is usually much smaller than the original length. Such dimension compression results in inevitable losses in discriminability and dimension robustness of the LLM embeddings, which motivates us to scale up the semantic representation. In this paper, we propose Mixture-of-Codes, which first constructs multiple independent codebooks for LLM representation in the indexing stage, and then utilizes the Semantic Representation along with a fusion module for the downstream recommendation stage. Extensive analysis and experiments demonstrate that our method achieves superior discriminability and dimension robustness scalability, leading to the best scale-up performance in recommendations.

推荐系统语义表示大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。