arXiv:2510.20817cs.LG2025-10被引 29

反向KL正则化会诱导模式坍缩,该研究提出改进方法提升模型生成多样性。

KL-Regularized Reinforcement Learning is Designed to Mode Collapse

  • 通过理论与实证揭示反向KL正则化本质导致单一模式输出
  • 低正则强度与等量奖励设置使目标分布趋向单峰,抑制多样性
  • 仅微调奖励尺度即可显著提升大模型与化学语言模型的解质量与多样性

通常认为优化反向KL会‘寻找模式’,而正向KL能‘覆盖质量’,后者更利于多模式采样。然而我们从数学和实验上表明,这一直觉在带反向/正向KL正则化的强化学习中并不成立。选择反向或正向KL实际上决定了最优目标分布族,其形式由正则化系数参数化。模式覆盖主要取决于正则强度及奖励与参考概率的相对尺度,而非正则化方向本身。进一步发现,常见设置如低正则强度、相等可验证奖励,会构造出单峰目标分布,即优化目标本质上非多样化。基于此,我们设计了一种简单、可扩展且理论合理的算法:仅对奖励幅度进行最小修改,但使目标分布对所有高质量采样模式赋予高概率。实验显示,该方法可有效后训练大语言模型与化学语言模型,在无外部多样性信号情况下提升解的质量与多样性,且适用于正向与反向KL,而原方法在两种情形下均失效。

原文摘要 · Abstract (English)

It is commonly believed that optimizing the reverse KL divergence results in "mode seeking", while optimizing forward KL results in "mass covering", with the latter being preferred if the goal is to sample from multiple diverse modes. We show -- mathematically and empirically -- that this intuition does not necessarily transfer well to doing reinforcement learning with reverse/forward KL regularization (e.g. as commonly used with language models). Instead, the choice of reverse/forward KL determines the family of optimal target distributions, parameterized by the regularization coefficient. Mode coverage depends primarily on other factors, such as regularization strength, and relative scales between rewards and reference probabilities. Further, we show commonly used settings such as low regularization strength and equal verifiable rewards tend to specify unimodal target distributions, meaning the optimization objective is, by construction, non-diverse. We leverage these insights to construct a simple, scalable, and theoretically justified algorithm. It makes minimal changes to reward magnitudes, yet optimizes for a target distribution which puts high probability over all high-quality sampling modes. In experiments, this simple modification works to post-train both Large Language Models and Chemical Language Models to have higher solution quality and diversity, without any external signals of diversity, and works with both forward and reverse KL when using either naively fails.

强化学习模式覆盖语言模型多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。