arXiv:2510.19178cs.LGcs.AI2025-10Conference of the …被引 16

强化学习微调中任务梯度不平衡,导致模型偏向某些任务。

Imbalanced Gradients in RL Post-Training of Multi-Task LLMs

  • 发现强化学习微调时不同任务梯度大小差异显著
  • 大梯度任务未必带来更好性能提升,存在学习收益错配
  • 适合关注多任务LLM训练公平性与优化策略的研究者

大型语言模型(LLMs)的多任务微调通常通过混合不同任务的数据并联合优化实现。该方法隐含假设所有任务贡献的梯度幅度相似;一旦这一假设不成立,优化将偏向梯度大的任务。本文揭示:在强化学习微调中,某些任务会产生显著更大的梯度,从而导致更新偏向这些任务。这种梯度不平衡本应合理,仅当大梯度对应大学习收益(即性能提升)时;但研究发现,大梯度任务的学习收益可能相似甚至更低。进一步分析表明,这种不平衡无法用典型训练统计量(如训练奖励或优势值)解释,说明其源于任务间的内在差异。这警示不应简单混合数据集,亟需未来工作探索针对梯度层面的规范化修正方法。

原文摘要 · Abstract (English)

Multi-task post-training of large language models (LLMs) is typically performed by mixing datasets from different tasks and optimizing them jointly. This approach implicitly assumes that all tasks contribute gradients of similar magnitudes; when this assumption fails, optimization becomes biased toward large-gradient tasks. In this paper, however, we show that this assumption fails in RL post-training: certain tasks produce significantly larger gradients, thus biasing updates toward those tasks. Such gradient imbalance would be justified only if larger gradients implied larger learning gains on the tasks (i.e., larger performance improvements) -- but we find this is not true. Large-gradient tasks can achieve similar or even much lower learning gains than small-gradient ones. Further analyses reveal that these gradient imbalances cannot be explained by typical training statistics such as training rewards or advantages, suggesting that they arise from the inherent differences between tasks. This cautions against naive dataset mixing and calls for future work on principled gradient-level corrections for LLMs.

强化学习多任务学习梯度平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。