arXiv:2504.21677cs.CL2025-04中稿 · SwissText 2025被引 2

构建了1.5万对瑞士新闻双语语料,用于跨语言研究。

20min-XD: A Comparable Corpus of Swiss News Articles

  • 基于语义相似性自动对齐法德双语新闻文章。
  • 覆盖2015至2024年,包含近似翻译到相关但不完全匹配的文本。
  • 适合跨语言模型训练与语言对比研究者使用。

我们提出20min-XD(20 Minuten cross-lingual document-level),一个来自瑞士在线新闻媒体20 Minuten/20 minutes的法德双语、文档级可比语料库。该数据集包含约15,000对新闻文章,时间跨度为2015至2024年,通过语义相似性实现自动对齐。本文详述了数据收集流程与对齐方法,并提供了语料库的定性与定量分析。结果表明,该语料库涵盖从近似翻译到松散相关文章的广泛跨语言相似性,适用于多种自然语言处理任务及语言学研究。我们已公开发布文档级和句子级对齐版本的数据集及实验代码。

原文摘要 · Abstract (English)

We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minuten/20 minutes. Our dataset comprises around 15,000 article pairs spanning 2015 to 2024, automatically aligned based on semantic similarity. We detail the data collection process and alignment methodology. Furthermore, we provide a qualitative and quantitative analysis of the corpus. The resulting dataset exhibits a broad spectrum of cross-lingual similarity, ranging from near-translations to loosely related articles, making it valuable for various NLP applications and broad linguistically motivated studies. We publicly release the dataset in document- and sentence-aligned versions and code for the described experiments.

双语语料跨语言新闻文本数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。