arXiv:2510.05125cs.CLcs.LG2025-10被引 4

让推荐系统理解商品ID就像说方言,提升精准度与解释性。

Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation

  • 将商品交互记录当作语言方言,用专家混合模型统一建模。
  • 在多个公开和私有数据集上表现优异,保持预训练模型文本理解力。
  • 适合需要自然语言查询与可解释推荐的场景。

协同过滤虽具预测精度与效率优势,大语言模型(LLM)则擅长表达与泛化推理,现代推荐系统亟需融合二者。用户对自然语言查询与透明解释的需求日益增长,推动统一方法的发展。然而,协同信号通常高效但语义模糊,而仅基于文本训练的LLM难以捕捉隐式用户偏好。本文提出项⽬-ID+口语混合专家语言模型(IDIOMoE),将物品交互历史视为语言空间中的原生方言,使协同信号能如自然语言般被理解。通过将预训练LLM每层的前馈网络拆分为独立的文本专家与物品专家,并使用令牌类型门控机制分离处理,该方法有效避免了文本与目录模态间的破坏性干扰。IDIOMoE在多个公开及私有数据集上均表现出色,同时保留了预训练模型的文本理解能力。

原文摘要 · Abstract (English)

While collaborative filtering delivers predictive accuracy and efficiency, and Large Language Models (LLMs) enable expressive and generalizable reasoning, modern recommendation systems must bring these strengths together. Growing user expectations, such as natural-language queries and transparent explanations, further highlight the need for a unified approach. However, doing so is nontrivial. Collaborative signals are often token-efficient but semantically opaque, while LLMs are semantically rich but struggle to model implicit user preferences when trained only on textual inputs. This paper introduces Item-ID + Oral-language Mixture-of-Experts Language Model (IDIOMoE), which treats item interaction histories as a native dialect within the language space, enabling collaborative signals to be understood in the same way as natural language. By splitting the Feed Forward Network of each block of a pretrained LLM into a separate text expert and an item expert with token-type gating, our method avoids destructive interference between text and catalog modalities. IDIOMoE demonstrates strong recommendation performance across both public and proprietary datasets, while preserving the text understanding of the pretrained model.

推荐系统大语言模型多模态融合对话推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。