arXiv:2504.11900cs.CL2025-04被引 23

用故事漏洞检测评估大模型的深层推理能力,发现其表现远不如人类。

Finding Flawed Fictions: Evaluating Complex Reasoning in Language Models via Plot Hole Detection

  • 通过可控算法生成故事漏洞,构建新评测基准
  • 顶尖大模型在长故事中漏洞识别准确率显著下降
  • 生成和摘要任务易引入漏洞,超过一半以上风险

故事是人类经验的核心。深入理解故事并发现情节漏洞——即违背故事内部逻辑或世界规则的不一致——需要复杂推理能力,包括实体与事件追踪、抽象思维、语用理解、常识与社会推理及心智理论。随着大语言模型(LLMs)越来越多地参与文本生成、解读与修改,严格评估其叙事一致性与深层语言理解至关重要。然而现有评测多聚焦表层理解。本文提出以故事漏洞检测为代理任务,评估LLMs的语言理解与推理能力。我们提出FlawedFictionsMaker算法,可可控且谨慎地在真人撰写的故事中合成漏洞。基于此,构建了鲁棒性强、经人工筛选保证高质量的评测基准FlawedFictions。实验发现,即使允许充分推理,当前最优LLMs在解决FlawedFictions任务上仍表现不佳,且随故事长度增加性能急剧下降。此外,基于LLM的摘要与生成任务极易引入新漏洞,相比原作,漏洞检出率分别提升50%以上与100%以上。

原文摘要 · Abstract (English)

Stories are a fundamental aspect of human experience. Engaging deeply with stories and spotting plot holes -- inconsistencies in a storyline that break the internal logic or rules of a story's world -- requires nuanced reasoning skills, including tracking entities and events and their interplay, abstract thinking, pragmatic narrative understanding, commonsense and social reasoning, and theory of mind. As Large Language Models (LLMs) increasingly generate, interpret, and modify text, rigorously assessing their narrative consistency and deeper language understanding becomes critical. However, existing benchmarks focus mainly on surface-level comprehension. In this work, we propose plot hole detection in stories as a proxy to evaluate language understanding and reasoning in LLMs. We introduce FlawedFictionsMaker, a novel algorithm to controllably and carefully synthesize plot holes in human-written stories. Using this algorithm, we construct a benchmark to evaluate LLMs' plot hole detection abilities in stories -- FlawedFictions -- , which is robust to contamination, with human filtering ensuring high quality. We find that state-of-the-art LLMs struggle in accurately solving FlawedFictions regardless of the reasoning effort allowed, with performance significantly degrading as story length increases. Finally, we show that LLM-based story summarization and story generation are prone to introducing plot holes, with more than 50% and 100% increases in plot hole detection rates with respect to human-written originals.

大模型评估叙事理解推理能力漏洞检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。