arXiv:2409.02076cs.CL2024-09ICLR被引 55

评测大模型生成长文本的能力,发现现有模型表现不佳。

LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs

  • 设计新基准测试长文本生成,包含四种场景与复杂指令。
  • 十款主流模型在32K长度下生成质量明显下降。
  • 适合关注长文本应用的开发者和研究者参考。

当前的基准测试如Needle-in-a-Haystack、Ruler和Needlebench侧重于模型理解长上下文输入的能力,但未能涵盖一个关键维度:生成高质量长文本。实际应用场景如设计方案、技术文档和创意写作需要在长序列中保持连贯性并遵循复杂指令,而现有基准无法充分评估这一能力。为此,我们提出LongGenBench,一个全新基准,用于严格评估大语言模型(LLMs)在遵循复杂指令的同时生成长文本的能力。该基准通过要求生成文本中包含特定事件或约束的任务,评估模型在四种不同场景、三种指令类型及两种生成长度(16K和32K tokens)下的表现。对十款先进LLMs的评估显示,尽管在Ruler测试中表现良好,所有模型在LongGenBench上均表现不佳,且随着文本长度增加,性能显著下降。这表明当前大模型尚未具备满足真实世界长文本生成需求的能力。

原文摘要 · Abstract (English)

Current benchmarks like Needle-in-a-Haystack (NIAH), Ruler, and Needlebench focus on models' ability to understand long-context input sequences but fail to capture a critical dimension: the generation of high-quality long-form text. Applications such as design proposals, technical documentation, and creative writing rely on coherent, instruction-following outputs over extended sequences - a challenge that existing benchmarks do not adequately address. To fill this gap, we introduce LongGenBench, a novel benchmark designed to rigorously evaluate large language models' (LLMs) ability to generate long text while adhering to complex instructions. Through tasks requiring specific events or constraints within generated text, LongGenBench evaluates model performance across four distinct scenarios, three instruction types, and two generation-lengths (16K and 32K tokens). Our evaluation of ten state-of-the-art LLMs reveals that, despite strong results on Ruler, all models struggled with long text generation on LongGenBench, particularly as text length increased. This suggests that current LLMs are not yet equipped to meet the demands of real-world, long-form text generation.

长文本生成模型评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。