提出新指标衡量大模型输出的原创性与质量平衡。
Measuring LLM Novelty As The Frontier Of Original And High-Quality Output
- 用未见词元比例与任务质量得分的调和平均定义新颖性度量。
- 模型规模扩大和后训练提升显著改善生成内容的新颖性。
- 提示工程对新颖性影响有限,常以牺牲质量为代价。
随着大语言模型在创意构思与科学发现中的应用日益广泛,评估其生成新颖内容的能力至关重要。以往研究将新颖性定义为相对于训练数据的原创性,但原创内容可能质量低下。而人类评分虽更可靠地衡量质量,却倾向于偏好记忆内容,限制了其作为度量标准的可靠性。本文提出一种新的新颖性度量:未见n-gram比例与任务特定质量得分的调和平均。基于此框架,我们分析了三种开源模型(OLMo、OLMo-2、Pythia)在故事续写、诗歌创作和创意工具使用三个任务上的表现。结果表明,部分基础模型生成的内容新颖性低于互联网上的人类写作。然而,增大模型规模和进行后训练能显著提升新颖性,主要得益于质量改善。在同一规模下,改进基础模型(如从OLMo 7B到OLMo-2 7B)也带来更高新颖性,源于更强的原创性。此外,推理时的提示策略(如提供新上下文示例)对新颖性影响较小,常在提升原创性的同时降低质量。这凸显了在创造性应用中亟需更有效的激发策略。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly used for ideation and scientific discovery, it is important to evaluate their ability to generate novel output. Prior work evaluates novelty as originality with respect to model training data, but original outputs may be of low quality. In contrast, non-expert judges more reliably score quality but may favor memorized outputs, limiting the reliability of human preference as a metric. We introduce a new novelty metric for LLM generations that balances originality and quality -- the harmonic mean of the fraction of \ngrams unseen during training and a task-specific quality score. Using this framework, we identify trends that affect the novelty of generations from three families of open-data models (OLMo, OLMo-2, and Pythia) on three creative tasks: story completion, poetry writing, and creative tool use. We find that model-generated text from some base LLMs is less novel than human-written text from the internet. However, increasing model scale and post-training reliably improves novelty due to improvements in output quality. We also find that improving the base model at the same scale (\eg OLMo 7B to OLMo-2 7B) leads to higher novelty due to higher originality. Finally, we observe that inference-time methods, such as prompting and providing novel in-context examples, have a much smaller effect on novelty, often increasing originality at the expense of quality. This highlights the need for further research into more effective elicitation strategies as we use models for creative applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。