arXiv:2603.05890cs.CLcs.AI2026-03ACL被引 7

发现大模型写长故事时会自相矛盾,提出新评测基准和检测工具。

Lost in Stories: Consistency Bugs in Long Story Generation by LLMs

  • 构建新基准ConStory-Bench,覆盖2000个提示和5类错误
  • 发现事实与时间错误最常见,多出现在故事中段和高熵文本区
  • 开发自动化检测工具,可定位具体矛盾并提供原文依据

大型语言模型虽能生成长达数万字的叙事,却常在情节、人物设定和世界规则上出现自相矛盾。现有评测侧重剧情质量和流畅性,忽视一致性问题。为此,我们提出ConStory-Bench,包含2000个提示和四种任务场景,定义五类错误及19种细分类别,并开发自动检测工具ConStory-Checker,通过显式文本证据判断矛盾。对多种LLM的评估显示:一致性错误在事实与时间维度最突出,多集中在故事中段,常出现在高熵文本片段,且某些错误类型易共现。这些发现为提升长篇叙事生成的一致性提供了方向。

原文摘要 · Abstract (English)

What happens when a storyteller forgets its own story? Large Language Models (LLMs) can now generate narratives spanning tens of thousands of words, but they often fail to maintain consistency throughout. When generating long-form narratives, these models can contradict their own established facts, character traits, and world rules. Existing story generation benchmarks focus mainly on plot quality and fluency, leaving consistency errors largely unexplored. To address this gap, we present ConStory-Bench, a benchmark designed to evaluate narrative consistency in long-form story generation. It contains 2,000 prompts across four task scenarios and defines a taxonomy of five error categories with 19 fine-grained subtypes. We also develop ConStory-Checker, an automated pipeline that detects contradictions and grounds each judgment in explicit textual evidence. Evaluating a range of LLMs through five research questions, we find that consistency errors show clear tendencies: they are most common in factual and temporal dimensions, tend to appear around the middle of narratives, occur in text segments with higher token-level entropy, and certain error types tend to co-occur. These findings can inform future efforts to improve consistency in long-form narrative generation. Our project page is available at https://picrew.github.io/constory-bench.github.io/.

故事生成一致性LLM评测错误分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。