arXiv:2608.26177cs.CLcs.AI2026-08

对比7种框架在长文本生成中的提纲表现,发现匹配度决定成败。

A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs

  • 构建统一评测基准,用AI评分直接评估提纲质量。
  • 不同框架在单章、多章、全书任务中表现各异,无绝对优劣。
  • 提纲评分与成文评分相关性弱,支持提纲与写作分离评估。

长文本生成暴露了大语言模型的根本局限:即使700亿参数模型在输出16,000词时也出现长度崩溃,多章节故事常出现‘中间迷失’的属性漂移现象。尽管‘先拟提纲再写作’模式已广泛应用,但现有研究仅评估最终成文,混淆了提纲与写作两个应分离的评价对象。本文构建了一个涵盖7种主流长文本生成框架、3种生成粒度(单章、多章、全书)的统一头对头评测基准,并提出基于锚点的LLM-as-a-judge协议,以5分制锚定评分体系直接评估提纲与源文本的一致性。在21个框架-粒度组合中,无单一框架全面领先,性能取决于框架输出形式与目标粒度的匹配度。SuperWriter在受限长度的单章模式下排名第一,但在全书模式中优势减弱。提纲评分与成文评分的相关性仅中等,支持提纲与写作解耦原则。受算力限制,写作侧评估仅覆盖部分案例;后续实验将扩大样本量并引入跨模型评估者,以增强统计推断能力。

原文摘要 · Abstract (English)

Long-form generation exposes fundamental limitations of large language models. Even 70B-parameter models exhibit length collapse at 16k-token outputs, and multi-chapter stories frequently trigger the attribute drift characteristic of the ``lost-in-the-middle'' effect. The ``outline-first, write-later'' paradigm has gained wide adoption, yet existing research evaluates the final writing rather than the outline itself, conflating two evaluation objects that should be decoupled. We construct a unified head-to-head benchmark covering 7 representative long-form generation frameworks across 3 generation granularities -- single-chapter, multi-chapter, and whole-book -- and propose an anchor-based LLM-as-a-judge protocol that directly assesses outlines against the source text on a 5-point anchored scale. Across 21 framework-granularity cells, no single framework dominates; performance depends on the match between a framework's intrinsic output form and the target granularity. SuperWriter ranks first in the length-constrained single-chapter mode, but this advantage degrades in whole-book mode. The outline-side ranking correlates only moderately with the writing-side ranking, supporting the outline--writing decoupling principle. Compute constraints limit the writing-side evaluation to a subset of cases; follow-up experiments will expand the sample size and add cross-model evaluators to enable stronger statistical inference.

长文本生成提纲评估LLM评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。