arXiv:2601.20546cs.CL2026-01Conference of the …被引 4

重新定义大模型创造力评估,提出更贴近人类认知的新方法。

Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models

  • 基于新颖性与合适性统一的理论,设计新评测任务CDAT
  • 小模型在创意上表现更好,大模型更重合适性而降低新颖性
  • 挑战传统评测方式,适合关注生成质量与真实创意的研究者

大语言模型在语言创造性任务中应用日益广泛,但现有评估体系缺乏对人类创造力理论的充分依据,难以解释。当前普遍采用的发散联想任务(DAT)仅关注新颖性,忽视了创造力核心要素——合适性。我们评估多个先进LLM在DAT上的表现,发现其得分反而低于两个无创造能力的基线,质疑DAT的有效性。基于人类创造力理论(新颖性+合适性),我们提出条件发散联想任务(CDAT),在保证上下文合适性的前提下评估新颖性,能更清晰区分噪声与真正创意,且方法简洁客观。结果显示,较小模型家族通常更具创意,而先进模型虽更注重合适性,但新颖性下降。我们推测训练与对齐过程可能使模型沿此权衡边界移动。相关数据集与代码已公开。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in verbal creative tasks. However, previous assessments of the creative capabilities of LLMs remain weakly grounded in human creativity theory and are thus hard to interpret. The widely used Divergent Association Task (DAT) focuses on novelty, ignoring appropriateness, a core component of creativity. We evaluate a range of state-of-the-art LLMs on DAT and show that their scores on the task are lower than those of two baselines that do not possess any creative abilities, undermining its validity for model evaluation. Grounded in human creativity theory, which defines creativity as the combination of novelty and appropriateness, we introduce Conditional Divergent Association Task (CDAT). CDAT evaluates novelty conditional on contextual appropriateness, separating noise from creativity better than DAT, while remaining simple and objective. Under CDAT, smaller model families often show the most creativity, whereas advanced families favor appropriateness at lower novelty. We hypothesize that training and alignment likely shift models along this frontier, making outputs more appropriate but less creative. We release the dataset and code.

创造力评估大模型测评生成质量人类认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。