让大模型生成更高质量且多样的内容,突破传统训练的单一输出瓶颈。
Jointly Reinforcing Diversity and Quality in Language Model Generations
- 用可学习的分区分割函数衡量深层语义多样性,超越表面词汇变化。
- 在五个创意任务上同时提升输出质量与新颖性,竞争数学任务中解法准确率和多样性双增。
- 首次证明主动优化多样性能促进强化学习中的探索,反而带来更高质量结果。
大语言模型后训练常以准确性和有用性为优先,导致输出分布变窄,降低创意与探索类任务(如头脑风暴、故事创作、问题求解)的实用性。本文提出多样性感知强化学习(DARLING),联合优化响应质量与语义多样性。核心是引入可学习的分区分割函数,度量超出表层词汇差异的深层多样性。该多样性信号与质量奖励结合,在在线强化学习中引导模型生成既优质又独特的输出。跨多个模型家族与规模的实验表明,DARLING在两类任务中均有效:非验证任务(指令遵循、创意写作)和可验证任务(竞赛数学)。在前一类任务的五个基准上,其输出在质量和新颖性上均优于仅优化质量的基线;在后一类任务中,达到更高的 pass@1(解题质量)和 pass@k(解法多样性)。最显著的是,显式优化多样性促进了在线强化学习中的探索行为,从而提升整体响应质量。
原文摘要 · Abstract (English)
Post-training of Large Language Models (LMs) often prioritizes accuracy and helpfulness at the expense of diversity. This creates a tension: while post-training improves response quality, it also sharpens output distributions and reduces the range of ideas, limiting the usefulness of LMs in creative and exploratory tasks such as brainstorming, storytelling, or problem solving. We address this challenge with Diversity-Aware Reinforcement Learning (DARLING), a framework that jointly optimizes for response quality and semantic diversity. At its core, DARLING introduces a learned partition function to measure diversity beyond surface-level lexical variations. This diversity signal is then combined with a quality reward during online reinforcement learning, encouraging models to generate outputs that are both high-quality and distinct. Experiments across multiple model families and sizes show that DARLING generalizes to two regimes: non-verifiable tasks (instruction following and creative writing) and verifiable tasks (competition math). On five benchmarks in the first setting, DARLING consistently outperforms quality-only RL baselines, producing outputs that are simultaneously of higher quality and novelty. In the second setting, DARLING achieves higher pass@1 (solution quality) and pass@k (solution variety). Most strikingly, explicitly optimizing for diversity catalyzes exploration in online RL, which manifests itself as higher-quality responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。