19个小型模型在新闻摘要上表现差异大,部分媲美70B大模型。
Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance
- 评估19个小型模型在2000条新闻上的摘要能力,关注相关性、连贯性等指标。
- 顶尖小模型如Phi3-Mini生成摘要更短,质量接近70B大模型。
- 简单提示词效果更好,指令微调对小模型提升不明显,适合资源受限场景。
资源受限环境下对高效摘要工具的需求日益增长。尽管大语言模型(LLMs)提供更优的摘要质量,但其高计算开销限制了实际应用。相比之下,小型语言模型(SLMs)更具可及性,可在边缘设备实现实时摘要。然而,其摘要能力及与LLMs的对比仍缺乏系统研究。本文对19个SLMs在2000条新闻样本上进行了全面评估,重点关注相关性、连贯性、事实一致性及摘要长度。结果表明,SLM性能差异显著,顶级模型如Phi3-Mini和Llama3.2-3B-Ins的性能可媲美70B LLMs,且生成摘要更短。值得注意的是,SLMs更适合简单提示词,复杂提示可能导致质量下降;此外,指令微调并未持续提升其新闻摘要能力。本研究不仅深化了对SLMs的理解,也为追求性能与资源平衡的高效摘要方案提供了实用指导。
原文摘要 · Abstract (English)
The increasing demand for efficient summarization tools in resource-constrained environments highlights the need for effective solutions. While large language models (LLMs) deliver superior summarization quality, their high computational resource requirements limit practical use applications. In contrast, small language models (SLMs) present a more accessible alternative, capable of real-time summarization on edge devices. However, their summarization capabilities and comparative performance against LLMs remain underexplored. This paper addresses this gap by presenting a comprehensive evaluation of 19 SLMs for news summarization across 2,000 news samples, focusing on relevance, coherence, factual consistency, and summary length. Our findings reveal significant variations in SLM performance, with top-performing models such as Phi3-Mini and Llama3.2-3B-Ins achieving results comparable to those of 70B LLMs while generating more concise summaries. Notably, SLMs are better suited for simple prompts, as overly complex prompts may lead to a decline in summary quality. Additionally, our analysis indicates that instruction tuning does not consistently enhance the news summarization capabilities of SLMs. This research not only contributes to the understanding of SLMs but also provides practical insights for researchers seeking efficient summarization solutions that balance performance and resource use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。