arXiv:2510.07919cs.LG2025-10被引 1

用自适应探索的强化学习,让推荐系统更懂用户个性化需求。

GRADE: Personalized Multi-Task Fusion via Group-relative Reinforcement Learning with Adaptive Dirichlet Exploration

  • 通过组间相对比较优化权重,避免依赖人工调参。
  • 在百万级用户场景下,点击率、转化率等指标提升超1%。
  • 适合大规模推荐系统中需要动态平衡多目标的场景。

在现代推荐与搜索系统中,平衡多个目标对用户满意度至关重要,但现有多任务融合(MTF)方法依赖静态的人工调参权重,难以捕捉个体用户意图。尽管强化学习(RL)提供了个性化路径,但传统方法常因训练不稳定和稀疏奖励而失效。为此,我们提出组相对强化学习与自适应狄利克雷探索(GRADE),一种新颖且鲁棒的个性化多任务融合框架。GRADE采用无评判器的组相对策略优化(GRPO),通过评估候选权重组的相对表现实现稳定高效的策略学习。其核心创新包括:使用狄利克雷分布对权重空间进行有结构的探索;设计融合稀疏用户反馈、密集模型先验与规则约束的复合奖励函数,有效引导搜索。在日活超亿级应用的内嵌市场中部署,GRADE在严格的大规模A/B测试中显著优于基线:点击率(CTR)+0.595%,转化率(CVR)+1.193%,订单每千次展示量(OPM)+1.788%,总订单量+1.568%。凭借优异表现,GRADE已在快手市场搜索场景全面上线,服务数亿用户。

原文摘要 · Abstract (English)

Balancing multiple objectives is critical for user satisfaction in modern recommender and search systems, yet current Multi-Task Fusion (MTF) methods rely on static, manually-tuned weights that fail to capture individual user intent. While Reinforcement Learning (RL) offers a path to personalization, traditional approaches often falter due to training instability and the sparse rewards inherent in these large-scale systems. To address these limitations, we propose Group-relative Reinforcement learning with Adaptive Dirichlet Exploration (GRADE), a novel and robust framework for personalized multi-task fusion. GRADE leverages a critic-free, Group Relative Policy Optimization (GRPO) paradigm, enabling stable and efficient policy learning by evaluating the relative performance of candidate weight groups. Its core innovations include employing the Dirichlet distribution for principled and structured exploration of the weight space, and a composite reward function that combines sparse user feedback with dense model priors and rule-based constraints to guide the search effectively. Deployed in the in-app marketplace of an application with over hundreds of millions daily active users, GRADE significantly outperforms established baselines, achieving substantial gains in rigorous large-scale A/B tests: +0.595\% in CTR, +1.193\% in CVR, +1.788\% in OPM, and +1.568\% in total order volume. Following its strong performance, GRADE has been fully deployed in the marketplace search scenario of Kuaishou, serving hundreds of millions of users.

推荐系统强化学习多任务融合个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。