arXiv:2602.15894cs.CLcs.LG2026-02被引 1

让大模型生成更多样化内容,同时保证质量不下降。

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

  • 在质量不下降前提下,通过熵最大化提升生成多样性。
  • 理论推导出闭式解,确保优化过程最优且可证明。
  • 支持在线与离线训练,适合需要多样化输出的场景。

在众多大语言模型对齐应用中,用户既期望输出高质量,又希望内容具备充分多样性。然而现有方法常面临质量与多样性之间的根本性权衡:提升质量往往降低多样性,增加多样性则可能损害质量。本文提出质量约束熵最大化策略优化(QEMPO),一种在显式保持输出质量的前提下增强大模型生成多样性的新框架。该方法具有坚实的理论基础:我们推导出一个闭式解析解,可在质量约束条件下严格最大化熵(多样性的一种原则性度量),并保证在给定目标下的最优性。基于此解,QEMPO自然适用于在线与离线训练设置。实验表明,QEMPO在不牺牲质量的前提下持续提升输出多样性,且在多数情况下实现质量和多样性的双重提升,与理论预期一致。

原文摘要 · Abstract (English)

In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, existing methods often face a fundamental trade-off between these objectives: approaches that improve output quality tend to reduce diversity, while methods that increase diversity often do so at the expense of quality. In this work, we propose Quality-constrained Entropy Maximization Policy Optimization (QEMPO), a novel framework that enhances the diversity of LLM outputs while explicitly preserving output quality. QEMPO is grounded in a strong theoretical foundation: we derive a closed-form analytical solution that provably maximizes entropy-a principled measure of diversity-subject to a quality constraint, with guarantees on optimality under the defined objective. Leveraging this solution, QEMPO naturally supports both online and offline training settings. Empirical results demonstrate that QEMPO consistently improves output diversity without sacrificing quality, and in many cases yields gains in both dimensions compared to existing baselines, aligning with our theoretical guarantees.

大模型多样性策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。