arXiv:2412.03151cs.AI2024-12被引 16

大模型创造力接近人类,集体表现堪比数名人类协作。

Large Language Models show both individual and collective creativity comparable to humans

  • 用13项任务对比大模型与人类个体及群体的创造力。
  • 顶级模型(Claude/GPT-4)创造力达人类52百分位,发散思维强但写作弱。
  • 10次提问后,大模型集体创造力相当于8-10人,适合团队型创意任务。

人工智能迄今主要自动化常规任务,若大语言模型(LLMs)具备与人类相当的创造力,将如何重塑未来工作?本研究通过13项跨领域的创造性任务,全面评估LLMs的表现。基准测试显示,最佳模型(Claude和GPT-4)在人类个体中排名52百分位;总体上,大模型在发散思维与问题解决方面表现优异,但在创意写作上仍落后。当对模型提出10次问题时,其集体创造力相当于8-10名人类的协作水平;进一步增加响应数量时,每多生成两个模型输出,相当于额外一个真实人类的贡献。因此,在最优应用下,大模型未来可能与小型人类团队竞争。

原文摘要 · Abstract (English)

Artificial intelligence has, so far, largely automated routine tasks, but what does it mean for the future of work if Large Language Models (LLMs) show creativity comparable to humans? To measure the creativity of LLMs holistically, the current study uses 13 creative tasks spanning three domains. We benchmark the LLMs against individual humans, and also take a novel approach by comparing them to the collective creativity of groups of humans. We find that the best LLMs (Claude and GPT-4) rank in the 52nd percentile against humans, and overall LLMs excel in divergent thinking and problem solving but lag in creative writing. When questioned 10 times, an LLM's collective creativity is equivalent to 8-10 humans. When more responses are requested, two additional responses of LLMs equal one extra human. Ultimately, LLMs, when optimally applied, may compete with a small group of humans in the future of work.

大模型创造力人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。