arXiv:2603.23183cs.IRcs.AI2026-03KDD被引 11

用语义标识增强大模型推理,提升生成式推荐效果。

Reasoning over Semantic IDs Enhances Generative Recommendation

论文配图:Reasoning over Semantic IDs Enhances Generative Recommendation
图 1 · 摘自论文原文
  • 通过强化语义标识与语言的对齐,让大模型更好理解物品编码。
  • 在三个真实数据集上显著提升推荐准确率,且改善可解释性。
  • 无需人工标注推理过程,适合研究生成式推荐与大模型融合者。

生成式推荐近年借助预训练大语言模型,将序列推荐建模为统一标记空间中的自回归生成任务,其中每个物品由紧凑的离散标记序列(即语义标识,SIDs)表示。这种基于SIDs的框架实现了大规模物品库的高效解码,并为基于大语言模型的推荐系统提供了利用丰富世界知识的自然接口。同时,大语言模型在推理能力上的突破推动了推理增强型推荐的发展,但针对SIDs的有效推理仍鲜有探索且面临挑战:物品标记对大语言模型而言本无语义意义,且推荐导向的SIDs推理难以评估,高质量监督信号稀缺。为此,我们提出SIDReasoner,一种两阶段框架,通过加强SIDs与语言的对齐来激发可迁移的大语言模型推理能力,而非依赖大量特定推荐的推理轨迹。具体而言,首先在由更强教师模型合成的丰富以SIDs为中心的数据集上进行多任务训练,使物品标记在多样语义与行为上下文中获得语义基础;在此基础上,进一步通过目标驱动的强化优化,引导模型走向有效的推理路径,而无需显式推理标注。在三个真实数据集上的大量实验验证了该方法的有效性。除准确性提升外,结果还揭示了大推理模型在生成式推荐中的更广泛应用潜力,包括提升可解释性和跨域泛化能力。

原文摘要 · Abstract (English)

Recent advances in generative recommendation have leveraged pretrained LLMs by formulating sequential recommendation as autoregressive generation over a unified token space comprising language tokens and itemic identifiers, where each item is represented by a compact sequence of discrete tokens, namely Semantic IDs (SIDs). This SID-based formulation enables efficient decoding over large-scale item corpora and provides a natural interface for LLM-based recommenders to leverage rich world knowledge. Meanwhile, breakthroughs in LLM reasoning motivate reasoning-enhanced recommendation, yet effective reasoning over SIDs remains underexplored and challenging. Itemic tokens are not natively meaningful to LLMs; moreover, recommendation-oriented SID reasoning is hard to evaluate, making high-quality supervision scarce. To address these challenges, we propose SIDReasoner, a two-stage framework that elicits reasoning over SIDs by strengthening SID--language alignment to unlock transferable LLM reasoning, rather than relying on large amounts of recommendation-specific reasoning traces. Concretely, SIDReasoner first enhances SID-language alignment via multi-task training on an enriched SID-centered corpus synthesized by a stronger teacher model, grounding itemic tokens in diverse semantic and behavioral contexts. Building on this enhanced alignment, SIDReasoner further improves recommendation reasoning through outcome-driven reinforced optimization, which guides the model toward effective reasoning trajectories without requiring explicit reasoning annotations. Extensive experiments on three real-world datasets demonstrate the effectiveness of our reasoning-augmented SID-based generative recommendation. Beyond accuracy, the results highlight the broader potential of large reasoning models for generative recommendation, including improved interpretability and cross-domain generalization.

生成式推荐大模型推理语义标识强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。