arXiv:2501.18128cs.CLcs.AI2025-01被引 4

对比20个语言模型,发现小模型也能高效生成新闻摘要。

Unraveling the Capabilities of Language Models in News Summarization

  • 系统评测20个模型在零样本与少样本下的摘要能力。
  • GPT-3.5-Turbo和GPT-4表现最佳,部分开源模型接近其水平。
  • 参考摘要质量差会拖累模型表现,适合关注性价比的开发者。

针对近期多种语言模型的涌现及对自然语言处理任务(尤其是摘要)的持续需求,本文对20个近期语言模型进行了全面基准测试,重点关注小型模型在新闻摘要任务中的表现。研究系统评估了这些模型在不同写作风格下、三种不同数据集上的摘要能力,聚焦于零样本与少样本学习设置,并采用融合自动指标、人工评价和大模型评判的稳健评估方法。有趣的是,在少样本设置中加入示范示例并未提升模型性能,甚至在某些情况下导致生成摘要质量下降。这一问题主要源于用于参考的黄金摘要质量较差,对模型产生负面影响。此外,研究结果凸显了GPT-3.5-Turbo与GPT-4的卓越表现,二者普遍领先。但在公开模型中,Qwen1.5-7B、SOLAR-10.7B-Instruct-v1.0、Meta-Llama-3-8B和Zephyr-7B-Beta展现出显著潜力,表明它们可作为大型模型的有力替代方案,适用于新闻摘要任务。

原文摘要 · Abstract (English)

Given the recent introduction of multiple language models and the ongoing demand for improved Natural Language Processing tasks, particularly summarization, this work provides a comprehensive benchmarking of 20 recent language models, focusing on smaller ones for the news summarization task. In this work, we systematically test the capabilities and effectiveness of these models in summarizing news article texts which are written in different styles and presented in three distinct datasets. Specifically, we focus in this study on zero-shot and few-shot learning settings and we apply a robust evaluation methodology that combines different evaluation concepts including automatic metrics, human evaluation, and LLM-as-a-judge. Interestingly, including demonstration examples in the few-shot learning setting did not enhance models' performance and, in some cases, even led to worse quality of the generated summaries. This issue arises mainly due to the poor quality of the gold summaries that have been used as reference summaries, which negatively impacts the models' performance. Furthermore, our study's results highlight the exceptional performance of GPT-3.5-Turbo and GPT-4, which generally dominate due to their advanced capabilities. However, among the public models evaluated, certain models such as Qwen1.5-7B, SOLAR-10.7B-Instruct-v1.0, Meta-Llama-3-8B and Zephyr-7B-Beta demonstrated promising results. These models showed significant potential, positioning them as competitive alternatives to large models for the task of news summarization.

新闻摘要语言模型小模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。