评测发现现有方法无法真实反映人类对AI创作的创意判断。
The Limits of Automatic Evaluation of Creativity in Large Language Models

- 用11维标准对比人类与自动评估对故事创意的打分
- 自动指标与人类评价相关性接近零,偏差严重
- 大模型裁判偏爱AI文本,误判人类创作的惊喜感
大型语言模型在生成需要创造力的文本方面已接近甚至超越人类表现,但如何评估其创造性仍是一大挑战。本文在WritingPrompts数据集上收集了人类对人类与AI生成短篇故事在11个维度上的创意评价,并将其与自动化客观指标及大模型作为评判者(LLM-as-a-Judge)的结果进行比较。实验显示,自动评估与人类判断存在显著偏差。特别是,基于大模型的裁判系统表现出对AI生成文本的系统性偏好,更倾向于其风格特征而非人类文本所体现的不可预测性等特质。相关性分析表明,广泛使用的自动指标在人类与AI生成内容上均与人类评价近乎无关,说明它们未能捕捉创造力的关键维度。研究揭示了当前自动评估创意文本方法的根本局限性,凸显将多维且主观的创造力简化为计算指标的困难。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity. We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, and compare these judgments with automated objective metrics and LLM-as-a-Judge evaluations. Our experiments reveal substantial misalignment between automatic evaluations and human assessments. In particular, LLM-based judges exhibit a systematic preference for AI-generated stories, consistently favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Furthermore, correlation analyses show that widely used automatic metrics exhibit near-zero alignment with human judgments across both human- and AI-generated stories, suggesting that they fail to capture important dimensions of creativity. These findings highlight fundamental limitations in current approaches to the automatic evaluation of creative text and underscore the difficulty of reducing the multidimensional and subjective nature of creativity to computational metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。