arXiv:2509.22641cs.CLcs.AI2025-09被引 4

用专家评价验证:单纯看文本新颖度,会误判大量AI生成内容。

Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity

  • 通过8618条人工标注,对比新颖性与恰当性对创意的影响。
  • 91%高新颖性表达被专家认为不具创意,且开源模型越新越不靠谱。
  • 顶尖模型仍难识别无意义表达,但自评新颖度比传统方法更准。

n-gram新颖性广泛用于评估语言模型生成文本是否超出训练数据范围,近年也被当作衡量文本创造力的指标。然而,创造性理论指出该方法不足,因其未考虑创造力的双重属性:新颖性(文本原创程度)与恰当性(语义合理性和实用性)。我们通过8,618条专家对人类与AI生成文本的细致阅读标注,研究了创造力与n-gram新颖性的关系。结果发现,尽管n-gram新颖性与专家评价的创造力正相关,但约91%的高新颖性表达未被视作具有创意,警示不可仅依赖此指标。此外,在开源大模型中,更高n-gram新颖性反而伴随更低恰当性;在前沿闭源模型的探索性研究中,其生成创意表达的概率低于人类。利用本数据集,我们测试零样本、少样本和微调模型识别专家认为新颖(积极)或不恰当(消极)表达的能力。总体上,前沿大模型表现远超随机,但仍存在改进空间,尤其在识别不恰当表达方面表现不佳。我们进一步发现,大模型自评的新颖性评分在分布外数据集上比n-gram方法更贴近专家偏好。

原文摘要 · Abstract (English)

N-gram novelty is widely used to evaluate language models' ability to generate text outside of their training data. More recently, it has also been adopted as a metric for measuring textual creativity. However, theoretical work on creativity suggests that this approach may be inadequate, as it does not account for creativity's dual nature: novelty (how original the text is) and appropriateness (how sensical and pragmatic it is). We investigate the relationship between this notion of creativity and n-gram novelty through 8,618 expert writer annotations of novelty, pragmaticality, and sensicality via close reading of human- and AI-generated text. We find that while n-gram novelty is positively associated with expert writer-judged creativity, approximately 91% of top-quartile n-gram novel expressions are not judged as creative, cautioning against relying on n-gram novelty alone. Furthermore, unlike in human-written text, higher n-gram novelty in open-source LLMs correlates with lower pragmaticality. In an exploratory study with frontier closed-source models, we additionally confirm that they are less likely to produce creative expressions than humans. Using our dataset, we test whether zero-shot, few-shot, and finetuned models are able to identify expressions perceived as novel by experts (a positive aspect of writing) or non-pragmatic (a negative aspect). Overall, frontier LLMs exhibit performance much higher than random but leave room for improvement, especially struggling to identify non-pragmatic expressions. We further find that LLM-as-a-Judge novelty ratings align with expert writer preferences in an out-of-distribution dataset, more so than an n-gram based metric.

文本生成创造力评估大模型评测语义恰当性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。