arXiv:2411.02316cs.CLcs.AI2024-11中稿 · ICCC 2025被引 33

对比60个AI与60人写短故事,发现AI文采好但创意不足。

Evaluating Creative Short Story Generation in Humans and Large Language Models

  • 用五句话线索生成任务测试人类与大模型的创意写作能力
  • AI故事语言复杂但新颖性、意外性和多样性不如普通人
  • 专家评价与自动评分一致,但非专家和AI高估了AI的创意

叙事是人类想象力的核心,依赖创造力生成新颖、有效且出人意料的故事。尽管大语言模型(LLMs)已展现出生成高质量故事的能力,其创造性写作潜力仍待深入探索。本文通过五句线索式创作任务,系统评估60个LLMs与60名人类在短篇故事生成中的创造力表现。采用自动指标衡量故事的新颖性、意外性、多样性和语言复杂度,并收集非专家与专家人类评分及图灵测试分类结果。结果显示,LLMs生成的故事在风格上更复杂,但在新颖性、意外性和多样性方面普遍弱于普通人类作者。专家评分与自动指标基本一致。然而,非专家及大模型自身均认为生成的文本更具创造性。我们探讨了这一评价差异的原因及其对人工与人类创造力研究的启示。

原文摘要 · Abstract (English)

Story-writing is a fundamental aspect of human imagination, relying heavily on creativity to produce narratives that are novel, effective, and surprising. While large language models (LLMs) have demonstrated the ability to generate high-quality stories, their creative story-writing capabilities remain under-explored. In this work, we conduct a systematic analysis of creativity in short story generation across 60 LLMs and 60 people using a five-sentence cue-word-based creative story-writing task. We use measures to automatically evaluate model- and human-generated stories across several dimensions of creativity, including novelty, surprise, diversity, and linguistic complexity. We also collect creativity ratings and Turing Test classifications from non-expert and expert human raters and LLMs. Automated metrics show that LLMs generate stylistically complex stories, but tend to fall short in terms of novelty, surprise and diversity when compared to average human writers. Expert ratings generally coincide with automated metrics. However, LLMs and non-experts rate LLM stories to be more creative than human-generated stories. We discuss why and how these differences in ratings occur, and their implications for both human and artificial creativity.

创意生成大模型人类对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。