arXiv:2608.15654cs.CLcs.AI2026-08

评测大模型在动态世界中持续讲好故事的能力,发现三要素难以兼得。

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

论文配图:When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations
图 1 · 摘自论文原文
  • 构建过程评估基准,分三项测持续生成、情节一致、发展丰富性。
  • 模型规模提升只改善持续生成,难提升情节一致与有意义发展。
  • 三能力相互竞争,现有加权方法无法选出最优配置,适合研究叙事智能者。

大语言模型能写出流畅故事,但开放式叙事需超越局部流畅性。在动态世界模拟与AI原生游戏中,模型必须维持事实、关系、因果依赖和角色状态随世界演变。我们提出WSE-bench,一个过程基准,分别评估持续生成、正典一致性与有意义发展的能力。生成覆盖率记录计划叙事步骤的完成比例;一致性追踪正典断裂情况;丰富性衡量玩家引导轨迹的有意义分支程度。在前沿模型中,一致性和丰富性不存在平滑权衡:其经验帕累托前沿非凹,存在多个非占优中间配置,无法通过正线性加权选择。额外结构可丰富轨迹,但不统一提升一致性,甚至可能缩短生成。模型规模主要提升持续生成能力,未带来正典一致性或有意义发展的可靠提升。结果表明,持续生成、正典一致性和有意义发展是独立且有时冲突的能力。WSE-bench通过将叙事评估从成品扩展到生成过程,揭示了这些动态。

原文摘要 · Abstract (English)

Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.

大模型叙事生成动态世界评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。