对比4大模型在5个数据集上的摘要能力,发现不同模型各有优劣。
Evaluating LLMs and Pre-trained Models for Text Summarization Across Diverse Datasets
- 用5个数据集测试BART、FLAN-T5等4个开源大模型的摘要性能
- 在CNN/DM和XSum上LLaMA-3-8B表现最优,但其他数据集表现不一
- 适合评估摘要模型选型或研究者参考基准结果
文本摘要在自然语言处理中至关重要,能将大量文本浓缩为简洁连贯的摘要。随着数字内容快速增长及信息检索需求上升,文本摘要成为近年研究热点。本研究系统评估了四种领先的预训练与开源大语言模型:BART、FLAN-T5、LLaMA-3-8B 和 Gemma-7B,涵盖五个多样化的数据集:CNN/DM、Gigaword、News Summary、XSum 以及 BBC News。采用广泛认可的自动评估指标(包括 ROUGE-1、ROUGE-2、ROUGE-L、BERTScore 与 METEOR)衡量模型生成连贯且信息丰富的摘要能力。结果揭示了各模型在处理不同类型文本时的相对优势与局限性。
原文摘要 · Abstract (English)
Text summarization plays a crucial role in natural language processing by condensing large volumes of text into concise and coherent summaries. As digital content continues to grow rapidly and the demand for effective information retrieval increases, text summarization has become a focal point of research in recent years. This study offers a thorough evaluation of four leading pre-trained and open-source large language models: BART, FLAN-T5, LLaMA-3-8B, and Gemma-7B, across five diverse datasets CNN/DM, Gigaword, News Summary, XSum, and BBC News. The evaluation employs widely recognized automatic metrics, including ROUGE-1, ROUGE-2, ROUGE-L, BERTScore, and METEOR, to assess the models' capabilities in generating coherent and informative summaries. The results reveal the comparative strengths and limitations of these models in processing various text types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。