arXiv:2604.09854cs.CL2026-04被引 1

用100种结局预测法量化故事悬念,让AI更懂人类叙事张力。

Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling

论文配图:Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling
图 1 · 摘自论文原文
  • 通过预测100个可能结局,评估每句话引发的悬念强度。
  • 新《纽约客》故事在该指标上远超AI生成作品,真实反映文学张力。
  • 结合叙事结构约束,提升故事悬念同时不牺牲现有评测表现。

当前大语言模型在创作故事方面表现参差不齐,且自身无法识别缺陷——在主流创意写作评测(EQ-Bench)中,零样本的AI故事被模型评委评为高于《纽约客》短篇小说,而后者是文学虚构作品的黄金标准。我们指出,现有评判体系忽略了引人入胜的故事情节的关键维度:叙事张力。为此,我们提出100-Endings度量方法:逐句推进故事,每一步让模型基于已有文本预测100种可能结局,以预测失败率衡量张力。除失败率外,句级曲线还衍生出曲率反转率(inflection rate),作为几何指标捕捉情节转折与揭示的频率。与传统评分方式不同,100-Endings能正确将《纽约客》故事排在所有AI输出之上。基于叙事学原理,我们设计了包含故事模板分析、思想构建与叙事骨架支撑的生成流水线。该方案显著提升了100-Endings指标下的叙事张力,同时保持在EQ-Bench排行榜上的性能水平。

原文摘要 · Abstract (English)

LLMs have so far failed both to generate consistently compelling stories and to recognize this failure--on the leading creative-writing benchmark (EQ-Bench), LLM judges rank zero-shot AI stories above New Yorker short stories, a gold standard for literary fiction. We argue that existing rubrics overlook a key dimension of compelling human stories: narrative tension. We introduce the 100-Endings metric, which walks through a story sentence by sentence: at each position, a model predicts how the story will end 100 times given only the text so far, and we measure tension as how often predictions fail to match the ground truth. Beyond the mismatch rate, the sentence-level curve yields complementary statistics, such as inflection rate, a geometric measure of how frequently the curve reverses direction, tracking twists and revelations. Unlike rubric-based judges, 100-Endings correctly ranks New Yorker stories far above LLM outputs. Grounded in narratological principles, we design a story-generation pipeline using structural constraints, including analysis of story templates, idea formulation, and narrative scaffolding. Our pipeline significantly increases narrative tension as measured by the 100-Endings metric, while maintaining performance on the EQ-Bench leaderboard.

叙事生成故事张力评测指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。