arXiv:2504.12522cs.CLcs.AI2025-04被引 55

提出衡量高质量内容多样性的新框架,发现调优模型实际更优。

Evaluating the Diversity and Quality of LLM Generated Content

  • 用质量阈值筛选输出,衡量真正有用的语义多样性
  • 偏好调优模型在高质量输出中多样性更高,反直觉但更实用
  • 小模型更高效生成独特内容,适合资源受限场景

近期研究指出,偏好调优技术(如基于人类反馈的强化学习PPO、GRPO及DPO)会降低模型输出多样性,造成部署困境。我们认为,脱离质量的多样性无实际价值。为此,提出测量有效语义多样性的框架——仅统计满足质量阈值的输出间的多样性。在无需人工干预的开放任务中,我们发现:若不考虑质量,偏好调优模型(尤其是强化学习训练)输出多样性较低;但当限定质量后,这些模型反而产生更高的有效语义多样性,优于监督微调(SFT)或基础模型。进一步分析显示,大模型虽有更高有效多样性,但小模型在固定采样预算下更参数高效地生成独特内容。该发现对需高质多样输出的应用(如创意辅助、合成数据生成)具有重要实践意义。

原文摘要 · Abstract (English)

Recent work suggests that preference-tuning techniques -- such as Reinforcement Learning from Human Feedback (RLHF) methods like PPO and GRPO, as well as alternatives like DPO -- reduce diversity, creating a dilemma given that these models are widely deployed in applications requiring varied outputs. We argue that diversity without consideration of quality has limited practical value. To address this issue, we introduce a framework for measuring effective semantic diversity -- diversity among outputs that meet quality thresholds -- which better reflects the practical utility of large language models (LLMs). Using open-ended tasks that require no human intervention, we find counterintuitive results: when using diversity metrics that do not explicitly consider quality, preference-tuned models -- particularly those trained via RL -- often produce outputs with lower diversity; however, these same preference-tuned models generate greater effective semantic diversity than supervised fine-tuned (SFT) or base models. Our analysis further shows another trend: while larger models may exhibit greater effective semantic diversity than smaller models, the smaller models are consistently more parameter-efficient at producing unique content within a fixed sampling budget. These findings have practical implications for applications that require diverse yet high-quality outputs, from creative assistance to synthetic data generation.

LLM评估多样性质量模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。