arXiv:2503.00847cs.CL2025-03EMNLP被引 7

用大模型提升论点摘要生成与评估,效果超越传统方法。

Argument Summarization and its Evaluation in the Era of Large Language Models

  • 用提示工程融合大模型改进摘要生成与评估流程。
  • 新系统在多个指标上达当前最优,Qwen-3-32B表现最佳。
  • 适合关注大模型应用与论点挖掘的研究者参考。

大型语言模型(LLMs)已革新自然语言生成任务,包括论点摘要(ArgSum),这是论点挖掘的重要分支。本文研究了先进LLMs在ArgSum系统中的集成及其评估方法。我们提出一种新型基于提示的评估方案,并通过一个全新的真人标注基准数据集进行验证。主要贡献包括:(i) 将LLMs融入现有ArgSum系统;(ii) 开发两种基于LLM的新ArgSum系统,并与先前方法对比;(iii) 引入先进的基于LLM的评估框架。实验表明,使用LLMs显著提升了摘要生成与评估性能,达到当前最优结果。在四个被测试的LLM中,参数最少的Qwen-3-32B表现最佳,甚至优于GPT-4o。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have revolutionized various Natural Language Generation (NLG) tasks, including Argument Summarization (ArgSum), a key subfield of Argument Mining. This paper investigates the integration of state-of-the-art LLMs into ArgSum systems and their evaluation. In particular, we propose a novel prompt-based evaluation scheme, and validate it through a novel human benchmark dataset. Our work makes three main contributions: (i) the integration of LLMs into existing ArgSum systems, (ii) the development of two new LLM-based ArgSum systems, benchmarked against prior methods, and (iii) the introduction of an advanced LLM-based evaluation scheme. We demonstrate that the use of LLMs substantially improves both the generation and evaluation of argument summaries, achieving state-of-the-art results and advancing the field of ArgSum. We also show that among the four LLMs integrated in (i) and (ii), Qwen-3-32B, despite having the fewest parameters, performs best, even surpassing GPT-4o.

论点摘要大模型评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。