arXiv:2606.06260cs.IRcs.AI2026-06被引 4

提出OneReason模型,让推荐系统像人一样思考。

OneReason Technical Report

论文配图:OneReason Technical Report
图 1 · 摘自论文原文
  • 预训练增强物品标记感知能力
  • 三层次认知增强的思维链格式
  • 专精再统一的强化学习训练法

OneRec系列生成式推荐模型已广泛部署于短视频、直播、广告和电商等实际服务中。然而,这些模型虽具备规模优势,其推理能力却难以激活,因无法构建仅由物品标记组成的有意义思维链(CoT)。受大语言模型中“先思考后回答”范式的启发,我们开展初步研究(如OneRec-Think、OpenOneRec)探索生成式推荐中的推理能力。但发现思考模式并未优于非思考模式。结合多模态语言模型中关于CoT鲁棒性的最新发现,我们认为有效推理依赖两个因素:感知(将物品标记锚定到语义层面)与认知(将用户行为序列重构为连贯的潜在兴趣点)。为此,我们提出OneReason,包含:(1) 预训练阶段强化物品标记感知;(2) 在监督微调中采用三级认知增强的CoT格式;(3) 在强化学习中采用专精再统一的训练策略以提升思考能力。

原文摘要 · Abstract (English)

Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic tokens only. Inspired by the success of the reasoning-style ``think before answer'' paradigm in the LLM field, we conduct preliminary studies (i.e., OneRec-Think, OpenOneRec) to explore reasoning capability in generative recommendation. Nevertheless, we notice an unexpected phenomenon: the thinking mode does not show advantages over the non-thinking mode. Drawing insights from recent findings on CoT robustness in multi-modal language models, we argue that effective reasoning in recommendation rests on two factors: perception, the ability to ground itemic tokens in their underlying language semantics, and cognition, the ability to reorganize a user's behavior sequence into coherent latent interest points. We therefore propose OneReason, which includes: (1) strong itemic token perception in pre-training, (2) a three-level cognition-enhanced CoT format for recommendation tasks in SFT, and (3) a specialize-then-unify training recipe in RL to enhance the thinking ability.

生成推荐思维链认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。