arXiv:2606.12790cs.CL2026-06

提出细粒度新颖性评估指标GENIE,精准衡量模型输出在任务中的创新程度。

GENIE: A Fine-Grained Measure for Novelty

论文配图:GENIE: A Fine-Grained Measure for Novelty
图 1 · 摘自论文原文
  • 基于任务特定特征,量化生成内容的新颖性
  • 相比整体指标,能揭示新颖性的多维特性
  • 适合研究模型创造力提升方法的效能

大语言模型在各项任务中持续表现出创造力和多样性不足。以往研究聚焦于模型是否具备生成创造性内容的能力,而本文关注新颖性,探究在特定任务下模型输出为何新颖或不新颖。我们提出细粒度评估指标GENIE,基于响应群体的特征,沿任务特定维度衡量输出的新颖性。结果表明,相较于GENIE,整体性指标难以捕捉新颖性的高维特性,且无法明确指向具体属性。最后,我们利用GENIE评估缓解创造力不足方法的有效性,以更深入理解这些方法在提升新颖性方面的潜力与局限。

原文摘要 · Abstract (English)

Large Language Models have consistently demonstrated a lack of creativity and diversity across tasks. Prior work has focused on addressing whether models are capable of generating creative outputs. Here, we aim to consider novelty and investigate what makes model-generated content novel or not novel in a task-specific manner. We propose a fine-grained evaluation metric GENIE to measure the novelty of responses along task-specific features with respect to a population of responses. We show that unlike GENIE, holistic metrics struggle to capture the high-dimensionality of novelty and do not provide insight on which properties they target. Finally, we use GENIE to measure the effectiveness of mitigation methods that address creativity to better understand where these methods can improve novelty.

新颖性评估大模型语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。