arXiv:2505.16245cs.CL2025-05EMNLP被引 9

解决大模型生成偏好中长度偏见,提升创意输出多样性。

Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models

  • 通过控制长度的筛选策略,平衡多样性、质量与长度。
  • 仅用3000组偏好数据显著提升词汇与语义多样性。
  • 小模型可作大模型的多样性导师,适合创意生成场景。

多样化的语言模型输出对创造性生成、开放任务和自我改进训练至关重要。我们发现,常见的多样性度量标准,甚至用于偏好优化的奖励模型,都会系统性地偏向较短输出,限制表达力。为此,我们提出一种长度可控的数据选择策略 Diverse-NS,可在保持长度一致的前提下提升响应多样性。通过生成并过滤兼顾多样性、质量与长度的偏好数据,Diverse-NS 仅需3000组偏好对即可实现有效训练。应用于 LLaMA-3.1-8B 和 Olmo-2 系列模型后,显著提升了词汇与语义多样性。在发散联想、人物设定、替代用途和创意写作四项任务中,多样性持续提升,响应质量仅轻微下降或略有增益。意外发现:在 Olmo-2 模型家族(7B 和 13B)中,较小的 Olmo-2-7B 可作为大模型的有效‘多样性教师’。通过显式消除长度偏见,该方法高效推动模型生成更丰富、更具表现力的输出。

原文摘要 · Abstract (English)

Diverse language model responses are crucial for creative generation, open-ended tasks, and self-improvement training. We show that common diversity metrics, and even reward models used for preference optimization, systematically bias models toward shorter outputs, limiting expressiveness. To address this, we introduce Diverse, not Short (Diverse-NS), a length-controlled data selection strategy that improves response diversity while maintaining length parity. By generating and filtering preference data that balances diversity, quality, and length, Diverse-NS enables effective training using only 3,000 preference pairs. Applied to LLaMA-3.1-8B and the Olmo-2 family, Diverse-NS substantially enhances lexical and semantic diversity. We show consistent improvement in diversity with minor reduction or gains in response quality on four creative generation tasks: Divergent Associations, Persona Generation, Alternate Uses, and Creative Writing. Surprisingly, experiments with the Olmo-2 model family (7B, and 13B) show that smaller models like Olmo-2-7B can serve as effective "diversity teachers" for larger models. By explicitly addressing length bias, our method efficiently pushes models toward more diverse and expressive outputs.

多样性提示工程长文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。