arXiv:2601.21452cs.LGcs.AI2026-01被引 1

提出SAGE优化器,解决生成推荐中冷启动与多样性问题

SAGE: Sequence-level Adaptive Gradient Evolution for Generative Recommendation

  • 通过序列级信号对齐和非对称自适应边界提升学习效率
  • 在Amazon和RecIF-Bench上实现冷启动召回率提升,多样性与精度同步改善
  • 适合需要高多样性与长尾推荐的生成式推荐系统研究者

基于强化学习的偏好优化正被广泛用于对齐列表级生成推荐系统与复杂的多目标用户反馈,但现有优化器如梯度有界策略优化(GBPO)在推荐场景中存在结构性缺陷。我们发现其存在对称保守性失效问题:对称更新约束抑制了稀有正信号(如冷启动商品)的学习,静态负样本约束无法防止以拒绝为主的反馈导致的多样性崩溃,组归一化多目标奖励则引发低分辨率训练信号。为此,我们提出SAGE(序列级自适应梯度演化)优化器,专为列表级生成推荐设计。SAGE引入基于几何均值重要性比率的序列级信号对齐,以及解耦的多目标优势估计器,降低令牌级方差并缓解奖励坍塌;同时采用非对称自适应边界机制,对成功推荐集实施正向增强,并结合熵感知惩罚抑制低多样性失败。在Amazon Product Reviews与大规模RecIF-Bench数据集上的实验表明,SAGE在顶K准确率、冷启动召回率和多样性方面均取得一致提升,且训练过程保持数值稳定。结果表明,非对称、序列感知的策略优化为解决生成推荐中的优化失败提供了原则性且有效的框架。

原文摘要 · Abstract (English)

Reinforcement learning-based preference optimization is increasingly used to align list-wise generative recommenders with complex, multi-objective user feedback, yet existing optimizers such as Gradient-Bounded Policy Optimization (GBPO) exhibit structural limitations in recommendation settings. We identify a Symmetric Conservatism failure mode in which symmetric update bounds suppress learning from rare positive signals (e.g., cold-start items), static negative-sample constraints fail to prevent diversity collapse under rejection-dominated feedback, and group-normalized multi-objective rewards lead to low-resolution training signals. To address these issues, we propose SAGE (Sequence-level Adaptive Gradient Evolution), a unified optimizer designed for list-wise generative recommendation. SAGE introduces sequence-level signal alignment via a geometric-mean importance ratio and a decoupled multi-objective advantage estimator to reduce token-level variance and mitigate reward collapse, together with asymmetric adaptive bounding that applies positive Boost updates to successful slates and an entropy-aware penalty to discourage low-diversity failures. Experiments on Amazon Product Reviews and the large-scale RecIF-Bench demonstrate consistent improvements in top-K accuracy, cold-start recall, and diversity across both Semantic-ID and native-text action spaces, while preserving numerical stability during training. These results suggest that asymmetric, sequence-aware policy optimization provides a principled and effective framework for addressing optimization failures in generative recommendation.

生成推荐强化学习冷启动多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。