不同大模型生成的创意高度相似,可能限制用户创造力。
We're Different, We're the Same: Creative Homogeneity Across LLMs
- 用标准化测试对比人类与多款大模型的创意输出
- 大模型间创意相似度远高于人与人之间
- 提示用户无论换哪个模型,都可能陷入相同创意圈
如今众多强大大语言模型被用作写作辅助、创意生成等工具。尽管它们被宣传为创意助手,但已有研究发现,与单一模型互动会缩小创意输出范围。然而这些研究仅关注单个模型,未回答这种局限是源于特定模型本身,还是使用大模型作为创意伙伴这一行为本身所致。为此,我们通过标准化创意测试获取人类及多种大模型的创意响应,并比较其群体层面的多样性。结果显示,大模型之间的创意相似度远高于人类彼此之间的相似度,即使在控制响应结构等关键变量后依然成立。这一发现揭示了当前大模型在创意输出上存在显著同质性,表明无论选用哪款大模型作为创意伙伴,都有可能使所有用户趋向于有限的一组‘创意’输出。
原文摘要 · Abstract (English)
Numerous powerful large language models (LLMs) are now available for use as writing support tools, idea generators, and beyond. Although these LLMs are marketed as helpful creative assistants, several works have shown that using an LLM as a creative partner results in a narrower set of creative outputs. However, these studies only consider the effects of interacting with a single LLM, begging the question of whether such narrowed creativity stems from using a particular LLM -- which arguably has a limited range of outputs -- or from using LLMs in general as creative assistants. To study this question, we elicit creative responses from humans and a broad set of LLMs using standardized creativity tests and compare the population-level diversity of responses. We find that LLM responses are much more similar to other LLM responses than human responses are to each other, even after controlling for response structure and other key variables. This finding of significant homogeneity in creative outputs across the LLMs we evaluate adds a new dimension to the ongoing conversation about creativity and LLMs. If today's LLMs behave similarly, using them as a creative partners -- regardless of the model used -- may drive all users towards a limited set of "creative" outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。