arXiv:2511.14598cs.CL2025-11Conference of the …

利用报纸标题页摘要自动生成低资源语言摘要数据集

Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages

  • 从报纸标题页的编辑摘要中提取自然生成的摘要数据
  • 在7种语言中验证该方法有效性,构建首个希伯来语多文档摘要数据集
  • 适用于不同资源水平的语言,推动低资源语言研究

高质量摘要数据在代表性不足的语言中仍然稀缺。然而,近年来数字化的历史报纸提供了大量未被利用的、自然标注的数据。本文提出一种新方法,通过报纸首页的‘预告摘要’(Front-Page Teasers)收集自然生成的摘要,这些摘要由编辑为长篇文章所作。我们发现这一现象在七种不同语言中普遍存在,并支持多文档摘要任务。为实现数据规模化,我们开发了适配不同语言资源水平的自动化流程。最后,将该流程应用于一份希伯来语报纸,构建出 HEBTEASESUM——首个专用于希伯来语的多文档摘要数据集。

原文摘要 · Abstract (English)

High quality summarization data remains scarce in under-represented languages. However, historical newspapers, made available through recent digitization efforts, offer an abundant source of untapped, naturally annotated data. In this work, we present a novel method for collecting naturally occurring summaries via Front-Page Teasers, where editors summarize full length articles. We show that this phenomenon is common across seven diverse languages and supports multi-document summarization. To scale data collection, we develop an automatic process, suited to varying linguistic resource levels. Finally, we apply this process to a Hebrew newspaper title, producing HEBTEASESUM, the first dedicated multi-document summarization dataset in Hebrew.

摘要数据低资源语言报纸数据多文档摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。