arXiv:2512.12839cs.CL2025-12ACL被引 9

首个长篇小说自动评估基准,揭示读者最关注的评价维度。

What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation

  • 构建600本新书的长故事评估基准,平均12.1万词,最大39.7万词。
  • 聚合与摘要式评估优于传统方法,尤其在细节捕捉和效率上表现突出。
  • 提出80亿参数模型NovelCritique,比GPT-4o更贴近人类评分。

本文针对超过10万词的长篇小说自动评估这一挑战性领域展开系统研究,聚焦两个核心问题:(1)读者最关注哪些评价维度;(2)有效的长篇故事评估方法。我们首次构建大规模基准LongStoryEval,包含600本新出版书籍,平均长度121K tokens(最大397K),每本书附带平均评分及按评价维度组织的读者评论。通过分析用户提及的评价维度,提出评估标准结构,并实验验证8个顶层维度中最具影响力者。对比三种评估方法——基于聚合、增量更新和基于摘要的评估,发现聚合与摘要方法更优,前者擅长细节评估,后者更具效率。基于此,提出8B规模模型NovelCritique,采用高效摘要框架,在指定维度上评审并打分故事,其结果在对齐人类评价方面超越GPT-4o。数据集与代码已开源。

原文摘要 · Abstract (English)

In this work, we conduct systematic research in a challenging area: the automatic evaluation of book-length stories (>100K tokens). Our study focuses on two key questions: (1) understanding which evaluation aspects matter most to readers, and (2) exploring effective methods for evaluating lengthy stories. We introduce the first large-scale benchmark, LongStoryEval, comprising 600 newly published books with an average length of 121K tokens (maximum 397K). Each book includes its average rating and multiple reader reviews, presented as critiques organized by evaluation aspects. By analyzing all user-mentioned aspects, we propose an evaluation criteria structure and conduct experiments to identify the most significant aspects among the 8 top-level criteria. For evaluation methods, we compare the effectiveness of three types: aggregation-based, incremental-updated, and summary-based evaluations. Our findings reveal that aggregation- and summary-based evaluations perform better, with the former excelling in detail assessment and the latter offering greater efficiency. Building on these insights, we further propose NovelCritique, an 8B model that leverages the efficient summary-based framework to review and score stories across specified aspects. NovelCritique outperforms commercial models like GPT-4o in aligning with human evaluations. Our datasets and codes are available at https://github.com/DingyiYang/LongStoryEval.

长文本评估小说生成自动评测LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。