arXiv:2506.10161cs.CLcs.AI2025-06被引 5

用叙事规划评估大模型写故事能力,发现其在因果逻辑上表现好,但角色动机和冲突设计仍难胜任。

Can LLMs Generate Good Stories? Insights and Challenges from a Narrative Planning Perspective

  • 用文学案例构建叙事规划评测基准,关注因果、动机、冲突三要素
  • GPT-4级模型小规模故事可做到因果合理,但角色意图与戏剧冲突仍不足
  • 适合研究生成式叙事、游戏剧情设计的学者与开发者参考

故事生成是大语言模型的重要应用之一。然而,由于自动评估方法受限及人工评估成本高、主观性强,对大模型生成高质量故事的能力理解仍不充分。计算叙事学为优质故事的构成提供了洞见,并被用于符号化叙事规划方法中。本文通过让大模型解决叙事规划问题,深入探究其故事生成能力。我们基于文学实例构建了一个评测基准,聚焦因果合理性、角色意图性与戏剧冲突。实验表明,GPT-4级别的模型在小规模故事中可生成因果连贯的内容,但在角色意图与戏剧冲突方面仍具挑战,需借助强化学习训练以提升复杂推理能力。结果揭示了大模型在不同维度维持故事质量的边界,同时展现出有趣的解题行为,为将大模型应用于游戏环境中的叙事规划提供了启示与思考。

原文摘要 · Abstract (English)

Story generation has been a prominent application of Large Language Models (LLMs). However, understanding LLMs' ability to produce high-quality stories remains limited due to challenges in automatic evaluation methods and the high cost and subjectivity of manual evaluation. Computational narratology offers valuable insights into what constitutes a good story, which has been applied in the symbolic narrative planning approach to story generation. This work aims to deepen the understanding of LLMs' story generation capabilities by using them to solve narrative planning problems. We present a benchmark for evaluating LLMs on narrative planning based on literature examples, focusing on causal soundness, character intentionality, and dramatic conflict. Our experiments show that GPT-4 tier LLMs can generate causally sound stories at small scales, but planning with character intentionality and dramatic conflict remains challenging, requiring LLMs trained with reinforcement learning for complex reasoning. The results offer insights on the scale of stories that LLMs can generate while maintaining quality from different aspects. Our findings also highlight interesting problem solving behaviors and shed lights on challenges and considerations for applying LLM narrative planning in game environments.

故事生成大模型叙事规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。