大模型生成的故事比人类更相似,缺乏多样性。
Do Large Language Models Always Tell The Same Stories?

- 用叙事相似性框架对比10个大模型与人类写作
- 模型故事相互间相似度高于人类作品,趋向统一模板
- 温度调节和负面提示等方法无法改善同质化问题
近年来大语言模型在生成高质量文本方面取得进展,但其输出多样性仍存争议。本文通过叙事相似性框架,基于 r/WritingPrompts 数据集的人类写作样本与提示,对10个代表性大模型进行叙事相似性评估,结合人工评价与三种自动标注方法。结果表明,模型生成的故事彼此间相似度显著高于人类作品。前沿模型尤其趋同于一种‘平均’通用叙事,虽能近似个体人类故事,却缺乏人类作者群体的多样性。此外,常见的缓解策略如负向提示和温度调节,均未能有效解决此同质化问题。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet the question of whether these models are capable of generating diverse outputs remains contested. In this work, we investigate the diversity of LLM-generated stories through the framework of narrative similarity. Using a contrastive framework and a dataset of human-written stories and prompts from r/WritingPrompts, we collect narrative similarity judgments across 10 representative LLMs, utilizing both human evaluations and three different automatic annotation methods. Our findings reveal a consistent trend: LLM-generated narratives are consistently more similar to each other than human-written stories are. We demonstrate that frontier models in particular converge on a ``mean'' generic narrative that approximates individual human stories but lacks the collective diversity of human authors. Finally, we show that common mitigation strategies, including negative prompting and temperature scaling, fail to meaningfully address this homogeneity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。