arXiv:2411.04585cs.CL2024-11NAACL综述被引 6

梳理133个摘要数据集,揭示领域研究碎片化问题

The State and Fate of Summarization Datasets: A Survey

  • 构建涵盖100多种语言的摘要数据集分类体系
  • 发现低资源语言高质量数据稀缺,新闻域依赖过重
  • 提供交互式网页与数据卡片模板,促进行业标准化

自动摘要因应用广泛而持续受关注。然而,我们发现标注工作分散且缺乏统一术语,导致难以发现现有资源或明确研究方向。为此,本文调研了涵盖133个数据集、超过100种语言的研究成果,创建了一个涵盖样本属性、收集方法和分布的新分类体系。基于该体系,我们发现低资源语言缺乏可访问的高质量数据集,且领域过度依赖新闻文本和自动获取的远距离监督数据。最后,我们发布了可交互的网页界面,用于探索分类体系与数据集集合,并提供摘要数据卡片模板,以推动未来研究走向更连贯的方向。

原文摘要 · Abstract (English)

Automatic summarization has consistently attracted attention due to its versatility and wide application in various downstream tasks. Despite its popularity, we find that annotation efforts have largely been disjointed, and have lacked common terminology. Consequently, it is challenging to discover existing resources or identify coherent research directions. To address this, we survey a large body of work spanning 133 datasets in over 100 languages, creating a novel ontology covering sample properties, collection methods and distribution. With this ontology we make key observations, including the lack in accessible high-quality datasets for low-resource languages, and the field's over-reliance on the news domain and on automatically collected distant supervision. Finally, we make available a web interface that allows users to interact and explore our ontology and dataset collection, as well as a template for a summarization data card, which can be used to streamline future research into a more coherent body of work.

数据集综述摘要生成语言多样性研究范式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。