构建长文本摘要与扩展的评测基准,支持多领域自动评估。
LCFO: Long Context and Long Form Output Dataset and Benchmarking
- 设计包含三类长度摘要的长文本数据集,实现摘要扩展可控评估。
- GPT-4o-mini在摘要和扩展任务中表现最佳,短摘要甚至超越人类。
- 发现现有自动指标与人工评分相关性弱,但对流畅性和引用准确性较准。
本文提出长上下文与长输出(LCFO)评测基准,用于评估跨多个领域的渐进式摘要与摘要扩展能力。该数据集包含平均5000词的长文档,每篇文档配有20%、10%和5%长度的三类摘要,以及约15个与内容相关的问答对。特别地,LCFO在7个领域中提供了问答对与摘要之间的对齐关系。其核心目标是建立从短文本生成长文本的可控框架,即摘要扩展。为构建评估体系,提供人工生成输出的人工评分,以及多种前沿大模型的结果。GPT-4o-mini在摘要与扩展任务中均取得最优自动系统得分,分别领先约+10%和+20%;在短摘要任务中甚至超过人类表现约+7%。总体自动指标与人工评分相关性较低(约0.4),但在流畅性和引用准确性等具体维度上相关性中等(约0.6)。
原文摘要 · Abstract (English)
This paper presents the Long Context and Form Output (LCFO) benchmark, a novel evaluation framework for assessing gradual summarization and summary expansion capabilities across diverse domains. LCFO consists of long input documents (5k words average length), each of which comes with three summaries of different lengths (20%, 10%, and 5% of the input text), as well as approximately 15 questions and answers (QA) related to the input content. Notably, LCFO also provides alignments between specific QA pairs and corresponding summaries in 7 domains. The primary motivation behind providing summaries of different lengths is to establish a controllable framework for generating long texts from shorter inputs, i.e. summary expansion. To establish an evaluation metric framework for summarization and summary expansion, we provide human evaluation scores for human-generated outputs, as well as results from various state-of-the-art large language models (LLMs). GPT-4o-mini achieves best human scores among automatic systems in both summarization and summary expansion tasks (~ +10% and +20%, respectively). It even surpasses human output quality in the case of short summaries (~ +7%). Overall automatic metrics achieve low correlations with human evaluation scores (~ 0.4) but moderate correlation on specific evaluation aspects such as fluency and attribution (~ 0.6).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。