arXiv:2412.02819cs.CLcs.AI2024-12ACL被引 2

构建中文长文本摘要基准,揭示大模型在长文总结中的真实表现

CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels

  • 基于中文小说构建多尺度摘要数据集,覆盖16k~128k字符长度
  • 发现大模型易生成主观评论,小模型反而更高效经济
  • 提出通过提示工程与微调优化长文本摘要,适合中文NLP研究者

大语言模型在长文本任务中已得到广泛研究,但长文本摘要数据集匮乏制约了该领域进展。为此,我们提出CNNSum,一个基于中文小说的多尺度长文本摘要基准,包含四类共695个经人工标注的样本,文本长度从16,000到128,000字符不等。我们对多种大语言模型进行评测,并开展详细的人工评估以分析异常输出类型。研究发现:(1) 高级大模型常生成大量主观评论,导致摘要模糊;(2) 当前长文本摘要主要依赖记忆能力,大模型优势难发挥,小模型更具成本效益;(3) 不同提示类型与模型版本组合导致性能差异显著,微调可缓解此问题,基础版模型表现更优;(4) 具有RoPE-base缩放的模型展现出强外推潜力,仅用短文本数据即可显著提升长文本摘要性能,但其他插值方法需谨慎选择;(5) 相较于其他基准,CNNSum提供更可靠的评估结果。我们已公开CNNSum数据集以推动后续研究(https://github.com/CxsGhost/CNNSum)。

原文摘要 · Abstract (English)

Large language models (LLMs) have been well-researched in various long-context tasks. However, the scarcity of long-context summarization datasets hinders progress in this area. To address this, we introduce CNNSum, a multi-scale long-context summarization benchmark based on Chinese novels, featuring human-driven annotations across four subsets totaling 695 samples, with lengths ranging from 16k to 128k. We benchmark numerous LLMs and conduct detailed human assessments to summarize abnormal output types. Furthermore, we extensively explore how to improve long-context summarization. In our study: (1) Advanced LLMs may generate much subjective commentary, leading to vague summaries. (2) Currently, long-context summarization mainly relies on memory ability. The advantages of Large LLMs are hard to utilize, thus small LLMs are more cost-effective. (3) Different prompt types paired with various version models may cause large performance gaps. In further fine-tuning, these can be mitigated, and the Base version models perform better. (4) LLMs with RoPE-base scaled exhibit strong extrapolation potential; using short-context data can significantly improve long-context summarization performance. However, further applying other interpolation methods requires careful selection. (5) CNNSum provides more reliable evaluation results than other benchmarks. We release CNNSum to advance future research.(https://github.com/CxsGhost/CNNSum)

长文本摘要中文NLP大模型评测数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。