arXiv:2502.19103cs.CL2025-02被引 20

提出长文本生成评估基准LongEval,揭示模型生成长文本时的性能退化规律。

LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm

  • 基于写作认知模型设计计划式生成范式,更贴近人类写作流程。
  • 发现模型规模虽影响生成能力,但小模型经充分训练后表现接近大模型。
  • 适合关注长文本生成、模型评估与写作辅助系统的研究者参考。

大型语言模型(LLMs)在多种自然语言处理任务中取得显著成功,但在长文本生成方面的能力仍不清晰且缺乏有效评估。我们的分析表明,当前的LLMs在应对长文本长度要求和信息密度时表现不佳,随着文本长度增加,性能明显下降。为定量定位这种性能退化并为模型发展提供进一步洞见,我们提出了LongEval,一个通过直接生成与计划式生成双范式评估长文本生成的基准,灵感来自认知与语言学中的写作模型。全面实验揭示有趣现象:尽管模型规模与生成能力相关,但经过长期文本充分训练的小模型(如LongWriter)表现可与大模型相当。所有代码与数据集已开源于https://github.com/Wusiwei0410/LongEval。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success in various natural language processing tasks, yet their ability to generate long-form content remains poorly understood and evaluated. Our analysis reveals that current LLMs struggle with length requirements and information density in long-text generation, with performance deteriorating as text length increases. To quantitively locate such a performance degradation and provide further insights on model development, we present LongEval, a benchmark that evaluates long-text generation through both direct and plan-based generation paradigms, inspired by cognitive and linguistic writing models. The comprehensive experiments in this work reveal interesting findings such as that while model size correlates with generation ability, the small-scale model (e.g., LongWriter), well-trained on long texts, has comparable performance. All code and datasets are released in https://github.com/Wusiwei0410/LongEval.

长文本生成模型评估计划式生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。