arXiv:2607.29010cs.IR2026-07

让推荐系统的推理过程自动进化,提升生成效率与准确性。

EvoReason: Self-Evolving Reasoning Primitive-Guided On-Policy Distillation for Latent Reasoning in Generative Recommendation

论文配图:EvoReason: Self-Evolving Reasoning Primitive-Guided On-Policy Distillation for Latent Reasoning in Generative Recommendation
图 1 · 摘自论文原文
  • 用可复用的推理模块构建结构化监督信号。
  • 通过闭环反馈优化学生模型的隐式推理能力。
  • 适合需要高效推理的生成式推荐系统研究者。

生成式推荐得益于增强推理的推理过程,而隐式推理通过将中间推理步骤编码为紧凑的连续表示,实现低延迟部署。然而,现有方法直接将原始思维链(CoT)轨迹蒸馏为隐式表示,忽略了推荐推理中冗余表达和不稳定路径的问题,导致监督信号不充分。为此,我们提出EvoReason,一种自演化隐式推理框架,通过基于推理原语的在线策略蒸馏,动态对齐显式推理监督与学生模型的隐式空间。首先,从高质量代理推荐轨迹中提取可复用的推理原语,每个原语捕捉核心推理行为,作为教师的伪工具。接着,赋予教师基于原语的推理能力,生成更少冗余、更一致的结构化CoT监督。最后,在隐式推理优化中引入自演化在线策略蒸馏机制,使原语引导的推理过程根据学生的隐式推理结果持续演化。通过这一闭环协同进化,策略更新不断改进隐式推理行为,实现更对齐的CoT监督与更有效的推理迁移。

原文摘要 · Abstract (English)

Generative recommendation benefits from reasoning-enhanced inference, and latent reasoning offers an efficient paradigm by encoding intermediate reasoning processes into compact continuous representations for latency-sensitive deployment. Despite its efficiency, existing latent reasoning approaches typically rely on directly distilling raw chain-of-thought (CoT) trajectories into latent representations, assuming that textual reasoning traces provide sufficient supervision. However, recommendation reasoning trajectories contain diverse reasoning processes with redundant expressions and unstable reasoning paths, making raw CoT supervision suboptimal for learning transferable latent reasoning representations. To address this challenge, we propose EvoReason, a self-evolving latent reasoning framework that adaptively aligns explicit reasoning supervision with the student's latent reasoning space through primitive-guided on-policy distillation. First, EvoReason extracts reusable reasoning primitives from high-quality agentic recommendation trajectories, where each primitive captures an essential reasoning behavior and serves as a pseudo-tool for structured teacher reasoning. Then, based on these primitives, we equip the teacher with primitive-aware reasoning capabilities, enabling it to generate structured CoT supervision with reduced redundancy and improved consistency. Finally, during latent reasoning optimization, EvoReason introduces a self-evolving on-policy distillation mechanism, where the primitive-guided reasoning process evolves according to the student's latent reasoning outcomes. Through this closed-loop co-evolution, policy updates continuously improve latent reasoning behaviors is refined according to the resulting latent reasoning outcomes, enabling progressively better-aligned CoT supervision and more effective reasoning transfer.

生成推荐隐式推理自演化知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。