arXiv:2608.19437cs.CLcs.AI2026-08

分析三年模型发现,大模型创意输出越来越趋同。

Are LLMs becoming similarly creative? Evidence from three years of models

论文配图:Are LLMs becoming similarly creative? Evidence from three years of models
图 1 · 摘自论文原文
  • 用嵌入相似度分析三类模型在开放任务中的创意表现
  • 发现模型输出多样性显著下降,创意趋于一致
  • 提醒警惕AI主导创作导致人类创造力被削弱

许多基准测试关注大语言模型(LLM)在可验证答案任务上的表现,但对其在开放式任务中——如创意、原创性和多样性——的表现演变了解较少。随着大模型越来越多地支持人类构思与创造性工作,理解其在开放式任务中的发展趋势至关重要。本文对三年间发布的多个模型在真实用户查询集Infinity-Chat100和经典心理学创造力测评任务Alternate Uses Task上的输出进行了初步分析。通过句子嵌入相似度衡量模型响应的差异性,结果表明,随着时间推移,模型输出的多样性出现统计显著下降,暗示不同模型在创意内容上逐渐趋同。若此趋势持续,基于大模型的同质化可能逐步削弱人类在人机共创中的自主性,亟需审慎评估大模型在人类创造性过程中的角色。

原文摘要 · Abstract (English)

Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly support human ideation and creative work, understanding trends in LLM performance on open-ended tasks is critical. This paper presents a preliminary analysis of LLM creative outputs spanning three years of model releases, examining model responses to Infinity-Chat100, a real-world collection of open-ended user queries, and the Alternate Uses Task, an established psychometric creativity assessment. Using sentence-embedding similarity, we examine trends in LLM responses to these prompts. Our findings show a statistically significant decrease in model output diversity over time, suggesting that LLM outputs may be converging in creative substance across models. If this trend persists, LLM-driven homogenization may progressively diminish human agency in human-AI co-creative work, demanding careful consideration of LLMs' role in the human creative process.

大模型创意生成多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。