对比测试发现,大模型在语言创造力上全面优于人类。
A Comparative Approach to Assessing Linguistic Creativity of Large Language Models and Humans
- 设计多任务测试评估人类与大模型的词汇创新能力
- 大模型在原创性、丰富性和灵活性上均胜过人类
- 人类偏重扩展式创造,大模型更倾向固定模式创造
本文提出一种面向人类和大型语言模型(LLMs)的通用语言创造力测试。测试包含多个任务,旨在评估其基于词法构词(派生与复合)及隐喻语言使用生成新词新短语的能力。我们对24名人类和同等数量的LLMs进行了测试,并利用OCSAI工具从原创性、丰富性和灵活性三个维度自动评估答案。结果表明,大模型在所有评估维度上均优于人类,且在八项任务中的六项表现更佳。进一步计算个体答案的独特性后发现,人类与大模型间存在细微差异。最后的简要人工分析显示,人类更倾向于E(扩展)型创造力,而大模型则偏好F(固定)型创造力。
原文摘要 · Abstract (English)
The following paper introduces a general linguistic creativity test for humans and Large Language Models (LLMs). The test consists of various tasks aimed at assessing their ability to generate new original words and phrases based on word formation processes (derivation and compounding) and on metaphorical language use. We administered the test to 24 humans and to an equal number of LLMs, and we automatically evaluated their answers using OCSAI tool for three criteria: Originality, Elaboration, and Flexibility. The results show that LLMs not only outperformed humans in all the assessed criteria, but did better in six out of the eight test tasks. We then computed the uniqueness of the individual answers, which showed some minor differences between humans and LLMs. Finally, we performed a short manual analysis of the dataset, which revealed that humans are more inclined towards E(extending)-creativity, while LLMs favor F(ixed)-creativity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。