arXiv:2508.19580cs.CL2025-08EMNLP被引 5

构建新数据集ArgCMV,推动大模型时代论点摘要研究

ArgCMV: An Argument Summarization Benchmark for the LLM-era

  • 基于真实在线辩论,用SOTA大模型构建12K条复杂论点
  • 相比旧数据集,论点更长、指代更多、主观表达更丰富
  • 揭示现有方法在新数据上表现下降,适合大模型摘要研究者

论点提取是论点摘要中的关键任务,旨在从论据中提取高层级简短总结。现有方法主要在流行的ArgKP21数据集上评估。本文指出ArgKP21数据集存在显著局限性,强调需建立更能代表真实人类对话的新基准。利用最先进的大语言模型(LLMs),我们构建了名为ArgCMV的新论点提取数据集,包含约12,000条来自超过3,000个主题的真实在线辩论论点。该数据集具有更高复杂性:论点更长、存在更多共指现象、主观话语单元占比更高、话题范围更广。实验表明,现有方法在ArgCMV上适应性差,并对现有基线和最新开源模型进行了广泛基准测试。本工作引入了一个面向长上下文在线讨论的新型论点提取数据集,为下一代大模型驱动的摘要研究奠定基础。

原文摘要 · Abstract (English)

Key point extraction is an important task in argument summarization which involves extracting high-level short summaries from arguments. Existing approaches for KP extraction have been mostly evaluated on the popular ArgKP21 dataset. In this paper, we highlight some of the major limitations of the ArgKP21 dataset and demonstrate the need for new benchmarks that are more representative of actual human conversations. Using SoTA large language models (LLMs), we curate a new argument key point extraction dataset called ArgCMV comprising of around 12K arguments from actual online human debates spread across over 3K topics. Our dataset exhibits higher complexity such as longer, co-referencing arguments, higher presence of subjective discourse units, and a larger range of topics over ArgKP21. We show that existing methods do not adapt well to ArgCMV and provide extensive benchmark results by experimenting with existing baselines and latest open source models. This work introduces a novel KP extraction dataset for long-context online discussions, setting the stage for the next generation of LLM-driven summarization research.

论点摘要大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。