利用报纸标题页摘要自动生成低资源语言摘要数据集
Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
- 从报纸标题页的编辑摘要中提取自然生成的摘要数据
- 在7种语言中验证该方法有效性,构建首个希伯来语多文档摘要数据集
- 适用于不同资源水平的语言,推动低资源语言研究
高质量摘要数据在代表性不足的语言中仍然稀缺。然而,近年来数字化的历史报纸提供了大量未被利用的、自然标注的数据。本文提出一种新方法,通过报纸首页的‘预告摘要’(Front-Page Teasers)收集自然生成的摘要,这些摘要由编辑为长篇文章所作。我们发现这一现象在七种不同语言中普遍存在,并支持多文档摘要任务。为实现数据规模化,我们开发了适配不同语言资源水平的自动化流程。最后,将该流程应用于一份希伯来语报纸,构建出 HEBTEASESUM——首个专用于希伯来语的多文档摘要数据集。
原文摘要 · Abstract (English)
High quality summarization data remains scarce in under-represented languages. However, historical newspapers, made available through recent digitization efforts, offer an abundant source of untapped, naturally annotated data. In this work, we present a novel method for collecting naturally occurring summaries via Front-Page Teasers, where editors summarize full length articles. We show that this phenomenon is common across seven diverse languages and supports multi-document summarization. To scale data collection, we develop an automatic process, suited to varying linguistic resource levels. Finally, we apply this process to a Hebrew newspaper title, producing HEBTEASESUM, the first dedicated multi-document summarization dataset in Hebrew.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。