arXiv:2508.05470cs.CL2025-08被引 14

四项创意评估方法在不同领域表现不一,揭示现有评测体系存在严重局限。

Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations

  • 对比四种创意评估方法在写作、解题和科研构思中的表现
  • 同一数据点上各方法结论常矛盾,跨领域一致性差
  • 适合关注创意评测可信度的研究者与模型开发者

我们考察并比较了四种代表性创意评估方法——困惑度、大模型作为评判者、创意指数(CI,衡量n-gram与网络语料重合度)、句法模板(检测常见词性模式重复)——在创意写作、非常规问题解决和研究思路生成等多元创意领域中的表现。针对每个领域,我们构建了包含人类标注的创意与非创意样本的数据集,并评估各指标区分两类样本的能力。分析发现,不同领域间及不同指标间均缺乏一致性:例如CI在创意写作中有效,但在问题解决中失效;同一数据点上,CI与困惑度甚至得出相反结论。我们指出主要局限:困惑度反映流畅性而非新颖性;大模型评判者对提示微小变化敏感且存在标签偏见;CI主要衡量词汇多样性,对实现细节高度敏感;句法模板在公式化语言环境中无效。结果强调需要更鲁棒、可泛化的评估框架,以更好对齐人类对创意的判断。

原文摘要 · Abstract (English)

We examine, analyze, and compare four representative creativity measures--perplexity, LLM-as-a-Judge, the Creativity Index (CI; measuring n-gram overlap with web corpora), and syntactic templates (detecting repetition of common part-of-speech patterns)--across the diverse creative domains, such as creative writing, unconventional problem-solving, and research ideation. For each domain, we compile datasets with human-aligned creative and uncreative examples and evaluate each metric's ability to discriminate between the two sets. Our analyses reveal limited consistency both across domains and metrics, as metrics that distinguish creativity in one domain fail in others (e.g., CI correctly distinguishes in creative writing but fails in problem-solving), and different metrics often disagree on the same data points (e.g., CI suggests one set to be more creative, while perplexity indicates the other set to be more creative.) We highlight key limitations, such as perplexity reflecting fluency rather than novelty; LLM-as-a-Judge producing inconsistent judgments under minor prompt variations and exhibiting bias towards particular labels; CI primarily measuring lexical diversity, with high sensitivity to implementation choices; and syntactic templates being ineffective in settings dominated by formulaic language. Our findings underscore the need for more robust, generalizable evaluation frameworks that better align with human judgments of creativity.

创意评估评测基准大模型评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。