量化大模型生成故事的剧情重复性,发现其创意多样性不足。
Echoes in AI: Quantifying lack of plot diversity in LLM outputs
- 提出Sui Generis分数,自动衡量剧情元素的独特性。
- 100个故事中,大模型生成内容常重复相同情节组合。
- 分数与人类对意外感的评价相关,适合评估创意质量。
随着大语言模型(LLMs)的快速发展,其在创意内容构思与生成中的应用日益广泛。一个关键问题浮现:当前的LLMs能否提供足够多样化的创意以真正促进集体创造力?我们研究了GPT-4和LLaMA-3两款先进模型在故事生成任务中的表现,发现生成的故事中经常出现跨多次生成、跨不同模型的剧情元素重复现象。为此,我们提出了Sui Generis分数,一种自动度量在相同提示下生成的多个故事情节中某一剧情元素独特性的指标。在100个短故事上的评估显示,大模型生成的内容频繁出现独特的剧情元素被反复复用,而真人原创故事的剧情几乎不会被复制或片段化重现。此外,人工评估表明,故事段落间Sui Generis分数的排序与人类对意外感的判断存在中等程度相关性,尽管该分数完全基于自动化计算,无需依赖人工标注。
原文摘要 · Abstract (English)
With rapid advances in large language models (LLMs), there has been an increasing application of LLMs in creative content ideation and generation. A critical question emerges: can current LLMs provide ideas that are diverse enough to truly bolster collective creativity? We examine two state-of-the-art LLMs, GPT-4 and LLaMA-3, on story generation and discover that LLM-generated stories often consist of plot elements that are echoed across a number of generations. To quantify this phenomenon, we introduce the Sui Generis score, an automatic metric that measures the uniqueness of a plot element among alternative storylines generated using the same prompt under an LLM. Evaluating on 100 short stories, we find that LLM-generated stories often contain combinations of idiosyncratic plot elements echoed frequently across generations and across different LLMs, while plots from the original human-written stories are rarely recreated or even echoed in pieces. Moreover, our human evaluation shows that the ranking of Sui Generis scores among story segments correlates moderately with human judgment of surprise level, even though score computation is completely automatic without relying on human judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。