arXiv:2510.18046cs.CLcs.AI2025-10

用大模型自动补全推荐系统的语义信息,提升推荐效果。

Language Models as Semantic Augmenters for Sequential Recommenders

  • 用少量样本提示大模型生成用户意图和物品关系的语义信号
  • 在多个基准任务上显著提升推荐模型性能,增强表示能力
  • 适合需要提升上下文理解的推荐系统研究者

大型语言模型(LLMs)擅长捕捉跨多种模态的潜在语义与上下文关系。但在基于序列交互数据建模用户行为时,若缺乏语义上下文,性能常受限。本文提出LaMAR框架,利用大模型在少样本设置下,通过现有元数据推断用户意图与物品间潜在语义关系,自动生成辅助上下文信号,如使用场景、物品意图或主题摘要,从而丰富原始序列的上下文深度。我们将这些生成信号融入基准序列建模任务中,结果表明其能持续提升性能。进一步分析显示,这些由大模型生成的信号具有高语义新颖性和多样性,增强了下游模型的表征能力。该工作开创了一种以数据为中心的新范式:大模型作为智能上下文生成器,提供一种半自动构建训练数据与语言资源的新方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) excel at capturing latent semantics and contextual relationships across diverse modalities. However, in modeling user behavior from sequential interaction data, performance often suffers when such semantic context is limited or absent. We introduce LaMAR, a LLM-driven semantic enrichment framework designed to enrich such sequences automatically. LaMAR leverages LLMs in a few-shot setting to generate auxiliary contextual signals by inferring latent semantic aspects of a user's intent and item relationships from existing metadata. These generated signals, such as inferred usage scenarios, item intents, or thematic summaries, augment the original sequences with greater contextual depth. We demonstrate the utility of this generated resource by integrating it into benchmark sequential modeling tasks, where it consistently improves performance. Further analysis shows that LLM-generated signals exhibit high semantic novelty and diversity, enhancing the representational capacity of the downstream models. This work represents a new data-centric paradigm where LLMs serve as intelligent context generators, contributing a new method for the semi-automatic creation of training data and language resources.

推荐系统大模型语义增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。