arXiv:2508.15308cs.IR2025-08被引 29

通过多路径推理与自我反思,提升推荐系统准确性与可靠性。

REG4Rec: Reasoning-Enhanced Generative Model for Large-Scale Recommendation Systems

  • 构建多动态语义路径,增强推荐多样性与表达能力。
  • 在多个真实数据集上优于基线模型,显著提升点击率与留存率。
  • 适合大规模推荐场景,尤其关注高精度与可解释性的研究者。

序列推荐旨在预测用户在大规模推荐系统中的下一步行为。传统方法常因信息交互不足而表现不佳,近年来的生成式推荐模型虽能直接生成物品预测,但受限于单一物品语义表示,存在推理路径单一、可靠性不强等问题。为此,我们提出REG4Rec,一种增强推理的生成式推荐模型,通过构建多条动态语义推理路径并引入自反思机制,确保高置信度推荐。具体地,REG4Rec采用基于MoE的并行量化码本(MPQ),为每个物品生成多个无序语义标记,拓展更大规模的多样化推理空间;同时设计训练阶段的推理增强策略,包括针对推荐任务的偏好对齐(PARS)和多步奖励增强(MSRA),以提升推理质量与泛化能力;推理阶段引入一致性导向的自反思剪枝(CORP),剔除不一致路径,防止错误传播。此外,我们开发了适用于大规模推荐的高效离线训练策略。实验证明,REG4Rec在真实数据集和线上评估中均表现优异,具备显著实际价值。

原文摘要 · Abstract (English)

Sequential recommendation aims to predict a user's next action in large-scale recommender systems. While traditional methods often suffer from insufficient information interaction, recent generative recommendation models partially address this issue by directly generating item predictions. To better capture user intents, recent studies have introduced a reasoning process into generative recommendation, significantly improving recommendation performance. However, these approaches are constrained by the singularity of item semantic representations, facing challenges such as limited diversity in reasoning pathways and insufficient reliability in the reasoning process. To tackle these issues, we introduce REG4Rec, a reasoning-enhanced generative model that constructs multiple dynamic semantic reasoning paths alongside a self-reflection process, ensuring high-confidence recommendations. Specifically, REG4Rec utilizes an MoE-based parallel quantization codebook (MPQ) to generate multiple unordered semantic tokens for each item, thereby constructing a larger-scale diverse reasoning space. Furthermore, to enhance the reliability of reasoning, we propose a training reasoning enhancement stage, which includes Preference Alignment for Reasoning (PARS) and a Multi-Step Reward Augmentation (MSRA) strategy. PARS uses reward functions tailored for recommendation to enhance reasoning and reflection, while MSRA introduces future multi-step actions to improve overall generalization. During inference, Consistency-Oriented Self-Reflection for Pruning (CORP) is proposed to discard inconsistent reasoning paths, preventing the propagation of erroneous reasoning. Lastly, we develop an efficient offline training strategy for large-scale recommendation. Experiments on real-world datasets and online evaluations show that REG4Rec delivers outstanding performance and substantial practical value.

推荐系统生成模型推理增强多路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。