大模型创造力未提升,同一模型输出差异极大。
Has the Creativity of Large-Language Models peaked? An analysis of inter- and intra-LLM variability
- 对比14个主流大模型在两类创意任务中的表现
- 仅0.28%生成内容达人类顶尖水平,且模型间性能波动大
- 提示词设计显著影响结果,需重复评估与精细调优
自2023年初ChatGPT广泛使用以来,多项研究声称大语言模型(LLMs)在创意任务中可媲美甚至超越人类。然而,其创造力是否随时间提升、输出是否稳定仍不明确。本研究评估了包括GPT-4、Claude、Llama、Grok、Mistral和DeepSeek在内的14个主流模型,在两个经验证的创意测评任务——发散联想任务(DAT)和替代用途任务(AUT)中的表现。结果显示,过去18至24个月间,模型在创意能力上无明显提升,且GPT-4的表现反而低于以往研究。在更常用的AUT任务中,所有模型平均优于普通人类,其中GPT-4o和o3-mini表现最佳。然而,仅有0.28%的模型生成内容达到人类前10%的创意基准。除模型间差异外,我们还发现显著的模型内变异性:同一模型在相同提示下,输出质量从低于平均到极具原创性不等。这种变异性对创意研究与实际应用具有重要影响。忽略该特性可能导致对模型创造力的误判,或高估或低估其能力。提示词的选择对不同模型影响各异。研究强调需建立更细致的评估框架,并重视模型选择、提示设计与重复测试,以在创造性场景中有效使用生成式AI工具。
原文摘要 · Abstract (English)
Following the widespread adoption of ChatGPT in early 2023, numerous studies reported that large language models (LLMs) can match or even surpass human performance in creative tasks. However, it remains unclear whether LLMs have become more creative over time, and how consistent their creative output is. In this study, we evaluated 14 widely used LLMs -- including GPT-4, Claude, Llama, Grok, Mistral, and DeepSeek -- across two validated creativity assessments: the Divergent Association Task (DAT) and the Alternative Uses Task (AUT). Contrary to expectations, we found no evidence of increased creative performance over the past 18-24 months, with GPT-4 performing worse than in previous studies. For the more widely used AUT, all models performed on average better than the average human, with GPT-4o and o3-mini performing best. However, only 0.28% of LLM-generated responses reached the top 10% of human creativity benchmarks. Beyond inter-model differences, we document substantial intra-model variability: the same LLM, given the same prompt, can produce outputs ranging from below-average to original. This variability has important implications for both creativity research and practical applications. Ignoring such variability risks misjudging the creative potential of LLMs, either inflating or underestimating their capabilities. The choice of prompts affected LLMs differently. Our findings underscore the need for more nuanced evaluation frameworks and highlight the importance of model selection, prompt design, and repeated assessment when using Generative AI (GenAI) tools in creative contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。