arXiv:2602.23639cs.IR2026-02被引 1

通过反思与修正提升生成式推荐质量,避免早期错误累积。

Learning to Reflect and Correct: Towards Better Decoding Trajectories for Large-Scale Generative Recommendation

  • 引入生成-反思-修正框架,分步优化推荐生成过程。
  • 在真实数据集上相比基线最高提升15.74%,广告收入增1.79%。
  • 适合工业级推荐系统优化,兼顾效果与推理效率。

生成式推荐(GR)已成为大规模推荐系统的有前景范式。然而,现有GR模型通常采用单次解码且无显式修正,导致早期偏差累积,最终降低推荐质量。为此,我们提出GRC,据知是首个面向GR的结构化反思-修正框架,将标准解码扩展为生成-反思-修正(GRC)流程。具体而言,GRC引入监督式反思-修正模板,将解码过程分解为初始草稿生成、多粒度反思和反思引导修正,从而在语义标记空间中实现结构化反思与修正。为进一步探索由GRC流程带来的扩大化修正空间,我们基于GRPO强化学习优化整个GRC轨迹,设计包含标记级与轨迹级信号的奖励函数。为支持高效在线服务,提出熵引导反思调度(EGRS)策略,在束搜索中动态为高不确定性轨迹分配更多修正预算。大量实验在真实数据集上显示,GRC持续优于六种先进基线,最高提升15.74%;线上A/B测试表明其在大规模工业推荐中具有显著价值,广告收入提升1.79%,仅带来适度延迟增加。

原文摘要 · Abstract (English)

Generative Recommendation (GR) has become a promising paradigm for large-scale recommendation systems. However, existing GR models typically perform single-pass decoding without explicit refinement, causing early deviations to accumulate and ultimately degrade recommendation quality. To tackle this problem, we propose GRC, which is, to our knowledge, the first structured reflection-correction framework for GR that extends standard decoding into a Generation-Reflection-Correction (GRC) process. Concretely, GRC introduces a supervised reflection-correction template that decomposes the decoding process into initial draft generation, multi-granular reflection, and reflection-guided correction, thereby enabling structured reflection and correction in the semantic token space. To further explore the enlarged refinement space introduced by the GRC process, we optimize the entire GRC trajectory with GRPO-based reinforcement learning, under a carefully designed reward function with token-level and trajectory-level signals. For efficient online serving, we propose an Entropy-Guided Reflection Scheduling (EGRS) strategy that dynamically allocates more correction budget to high-uncertainty decoding trajectories during beam search. Extensive experiments on real-world datasets show that GRC consistently outperforms six state-of-the-art baselines by up to 15.74%, and online A/B tests demonstrate its substantial practical value in large-scale industrial recommendation, delivering a 1.79% lift in advertising revenue with only modest latency overhead.

生成推荐强化学习工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。