arXiv:2507.05123cs.CLcs.AI2025-07中稿 · publication in the…被引 10

用提示工程评估6大模型在4类文本上的摘要表现,发现分段策略能显著提升长文档效果。

An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques

  • 采用零样本和上下文学习提示技术,测试不同模型在多数据集的表现。
  • 长文档经分段处理后,科学文本摘要准确率明显提升,且推理时间可控。
  • 结果对选择模型、设计提示及优化效率有实际指导意义,适合做NLP应用开发。

大型语言模型(LLMs)在自然语言处理中展现出生成类人文本的强大能力。尽管成果显著,其在新闻、对话、科研等多领域文本摘要任务中的表现尚未全面评估。同时,如何在不依赖大量训练数据的前提下实现高效摘要,已成为关键瓶颈。为此,我们系统评估了六种主流LLMs在四个数据集上的表现:CNN/Daily Mail和NewsRoom(新闻)、SAMSum(对话)、ArXiv(科研)。通过零样本与上下文学习提示工程,结合ROUGE与BERTScore指标进行评价,并分析推理耗时,以权衡质量与效率。针对长文档,提出基于句子的分块策略,使上下文窗口有限的模型可分阶段处理。结果显示,模型在新闻与对话任务中表现良好,而在科研长文档上经分块处理后性能显著提升。不同模型参数、数据集特性和提示设计均导致明显性能差异。研究为指令式NLP系统的高效部署提供了可操作洞见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) continue to advance natural language processing with their ability to generate human-like text across a range of tasks. Despite the remarkable success of LLMs in Natural Language Processing (NLP), their performance in text summarization across various domains and datasets has not been comprehensively evaluated. At the same time, the ability to summarize text effectively without relying on extensive training data has become a crucial bottleneck. To address these issues, we present a systematic evaluation of six LLMs across four datasets: CNN/Daily Mail and NewsRoom (news), SAMSum (dialog), and ArXiv (scientific). By leveraging prompt engineering techniques including zero-shot and in-context learning, our study evaluates the performance using the ROUGE and BERTScore metrics. In addition, a detailed analysis of inference times is conducted to better understand the trade-off between summarization quality and computational efficiency. For Long documents, introduce a sentence-based chunking strategy that enables LLMs with shorter context windows to summarize extended inputs in multiple stages. The findings reveal that while LLMs perform competitively on news and dialog tasks, their performance on long scientific documents improves significantly when aided by chunking strategies. In addition, notable performance variations were observed based on model parameters, dataset properties, and prompt design. These results offer actionable insights into how different LLMs behave across task types, contributing to ongoing research in efficient, instruction-based NLP systems.

文本摘要提示工程大模型评估长文档处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。