arXiv:2410.09773cs.CL2024-10NAACL被引 8

构建首个多语言多文档新闻摘要数据集,推动跨语言新闻总结研究

A Mixed-Language Multi-Document News Summarization Dataset and a Graphs-Based Extract-Generate Model

  • 基于图结构的提取-生成模型,融合多语言文档信息
  • 涵盖4种语言、10992个文档簇与摘要对,规模领先
  • 开源数据集与代码,助力跨语言摘要算法研发

现有新闻摘要研究主要聚焦单语言单文档(SLSD)、单语言多文档(SLMD)或跨语言单文档(CLSD)。但在真实场景中,国际事件相关新闻常涉及多语言、多文档,即混合语言多文档(MLMD)情形。因此,MLMD新闻摘要具有重要意义。然而,该领域缺乏相应数据集,制约了研究进展。为此,本文构建了一个名为MLMD-news的多语言多文档新闻摘要数据集,包含四种语言,共10,992个源文档簇与目标摘要对。同时提出一种基于图的提取-生成模型,并在该数据集上基准测试多种方法,公开发布数据集与代码(https://github.com/Southnf9/MLMD-news),旨在推动MLMD场景下摘要技术的发展。

原文摘要 · Abstract (English)

Existing research on news summarization primarily focuses on single-language single-document (SLSD), single-language multi-document (SLMD) or cross-language single-document (CLSD). However, in real-world scenarios, news about a international event often involves multiple documents in different languages, i.e., mixed-language multi-document (MLMD). Therefore, summarizing MLMD news is of great significance. However, the lack of datasets for MLMD news summarization has constrained the development of research in this area. To fill this gap, we construct a mixed-language multi-document news summarization dataset (MLMD-news), which contains four different languages and 10,992 source document cluster and target summary pairs. Additionally, we propose a graph-based extract-generate model and benchmark various methods on the MLMD-news dataset and publicly release our dataset and code\footnote[1]{https://github.com/Southnf9/MLMD-news}, aiming to advance research in summarization within MLMD scenarios.

新闻摘要多语言图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。