arXiv:2504.05228cs.CL2025-04被引 68

测试大模型生成多样新奇内容的能力,发现越大越单一。

NoveltyBench: Evaluating Language Models for Humanlike Diversity

  • 用精心设计的提示和真实用户问题构建评测集
  • 20个主流模型多样性远低于人类,大模型反而更单一
  • 提示技巧可部分提升多样性,但本质缺乏分布多样性

语言模型在标准评测中表现优异,却面临模式崩溃问题,即难以生成多样且新颖的输出。本文提出 NoveltyBench,一个专门评估语言模型生成多类高质、独特输出能力的基准。该基准采用能激发多样化回答的提示和经筛选的真实用户查询。对20个领先语言模型的评估显示,当前最先进系统生成的多样性显著低于人类写作者。值得注意的是,同一模型族中,较大模型的多样性反而低于较小版本,挑战了标准评测能力与生成实用性直接相关的认知。尽管如上下文再生等提示策略可激发多样性,研究结果仍揭示当前模型存在根本性的分布多样性缺失,限制其在需要多样化响应场景中的实用性,并呼吁建立以多样性与质量并重的新训练与评估范式。

原文摘要 · Abstract (English)

Language models have demonstrated remarkable capabilities on standard benchmarks, yet they struggle increasingly from mode collapse, the inability to generate diverse and novel outputs. Our work introduces NoveltyBench, a benchmark specifically designed to evaluate the ability of language models to produce multiple distinct and high-quality outputs. NoveltyBench utilizes prompts curated to elicit diverse answers and filtered real-world user queries. Evaluating 20 leading language models, we find that current state-of-the-art systems generate significantly less diversity than human writers. Notably, larger models within a family often exhibit less diversity than their smaller counterparts, challenging the notion that capability on standard benchmarks translates directly to generative utility. While prompting strategies like in-context regeneration can elicit diversity, our findings highlight a fundamental lack of distributional diversity in current models, reducing their utility for users seeking varied responses and suggesting the need for new training and evaluation paradigms that prioritize diversity alongside quality.

语言模型多样性评估生成质量模式崩溃

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。