arXiv:2410.04199cs.CLcs.AI2024-10EMNLP被引 34

首个专测长文本生成能力的基准,揭示大模型在长文创作中的严重退化问题

LongGenBench: Long-context Generation Benchmark

  • 设计可调长上下文的合成测试集,要求模型输出连贯长文本答案
  • 多数模型在长文本生成中性能下降1.2%至47.1%,表现显著退化
  • 发现Gemini-1.5-Flash和Qwen2系列在开源与闭源模型中相对更稳定

当前长上下文评测多聚焦于检索类任务,如针在草堆(NIAH)基准,要求大语言模型在长输入中定位特定信息。而长上下文生成指模型在长篇幅文本中生成连贯且符合语境的输出能力。尽管近期模型在NIAH等检索类评测中表现优异,但缺乏对长文本生成能力的系统评估。为此,我们提出一个合成基准LongGenBench,支持自定义生成上下文长度。该基准通过重构问题格式,要求模型以单一连贯长回答作答,突破传统评测局限。在广泛评估中发现:(1) 无论是API访问还是开源模型,长文本生成性能普遍下降,降幅达1.2%至47.1%;(2) 不同模型系列表现差异明显,其中Gemini-1.5-Flash在闭源模型中退化最小,Qwen2系列在开源模型中表现最稳定。

原文摘要 · Abstract (English)

Current long-context benchmarks primarily focus on retrieval-based tests, requiring Large Language Models (LLMs) to locate specific information within extensive input contexts, such as the needle-in-a-haystack (NIAH) benchmark. Long-context generation refers to the ability of a language model to generate coherent and contextually accurate text that spans across lengthy passages or documents. While recent studies show strong performance on NIAH and other retrieval-based long-context benchmarks, there is a significant lack of benchmarks for evaluating long-context generation capabilities. To bridge this gap and offer a comprehensive assessment, we introduce a synthetic benchmark, LongGenBench, which allows for flexible configurations of customized generation context lengths. LongGenBench advances beyond traditional benchmarks by redesigning the format of questions and necessitating that LLMs respond with a single, cohesive long-context answer. Upon extensive evaluation using LongGenBench, we observe that: (1) both API accessed and open source models exhibit performance degradation in long-context generation scenarios, ranging from 1.2% to 47.1%; (2) different series of LLMs exhibit varying trends of performance degradation, with the Gemini-1.5-Flash model showing the least degradation among API accessed models, and the Qwen2 series exhibiting the least degradation in LongGenBench among open source models.

长文本生成模型评测大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。