arXiv:2412.11317cs.CL2024-12中稿 · COLING 2024被引 4

构建首个大规模罗马尼亚语新闻摘要数据集,支持多任务生成研究

RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation

  • 从罗/摩新闻网站爬取并清洗61.5万篇带摘要的新闻
  • 包含标题、关键词、方言等元信息,支持多任务训练
  • 为罗马尼亚语生成模型提供基准测试数据,适合语言资源匮乏场景

监督式自动摘要方法依赖于充足的文档-摘要配对语料库。与自然语言处理中多数任务类似,现有摘要数据集主要为英语,制约了其他语言摘要模型的发展。为此,本文提出RoLargeSum,一个从罗马尼亚及摩尔多瓦公开新闻网站爬取并经严格清洗的大规模罗马尼亚语摘要数据集。该数据集包含超过61.5万篇新闻文章,每篇均配有摘要、标题、关键词、方言及其他元数据。我们进一步评估了多种BART变体及开源大模型在该数据集上的表现,用于基准测试。通过人工评估最佳模型输出,分析数据集潜在缺陷及未来改进方向。

原文摘要 · Abstract (English)

Using supervised automatic summarisation methods requires sufficient corpora that include pairs of documents and their summaries. Similarly to many tasks in natural language processing, most of the datasets available for summarization are in English, posing challenges for developing summarization models in other languages. Thus, in this work, we introduce RoLargeSum, a novel large-scale summarization dataset for the Romanian language crawled from various publicly available news websites from Romania and the Republic of Moldova that were thoroughly cleaned to ensure a high-quality standard. RoLargeSum contains more than 615K news articles, together with their summaries, as well as their headlines, keywords, dialect, and other metadata that we found on the targeted websites. We further evaluated the performance of several BART variants and open-source large language models on RoLargeSum for benchmarking purposes. We manually evaluated the results of the best-performing system to gain insight into the potential pitfalls of this data set and future development.

新闻摘要多语言NLP数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。