arXiv:2508.16371cs.CL2025-08Conference of the …被引 3

首个罗曼什语方言平行语料库,助力小语种机器翻译

The Mediomatix Corpus: Parallel Data for Romansh Language Varieties via Comparable Schoolbooks

  • 基于291本内容可比教材,自动对齐生成多语言段落
  • 获207,000个平行段落,总词元超200万,人工评估高度平行
  • 适合小语种NLP研究者,尤其关注瑞士罗曼什语的翻译应用

瑞士罗曼什语的五个语域(即方言)已高度标准化,并在各自社区的学校中教授。本文首次构建了罗曼什语各方言的平行语料库。该语料库基于291本内容可比的教材,采用自动对齐方法提取出20.7万个多语言并行段落,总计超过200万词元。小规模人工评估确认这些段落具有高度平行性,适用于罗曼什语方言间的机器翻译等自然语言处理任务。我们以CC-BY-NC-SA许可发布语料库的平行与非对齐版本,并通过在该数据集上训练和评估大语言模型及监督式多语言机器翻译模型,验证了其实用性。

原文摘要 · Abstract (English)

The five idioms (i.e., varieties) of the Romansh language are largely standardized and are taught in the schools of the respective communities in Switzerland. In this paper, we present the first parallel corpus of Romansh idioms. The corpus is based on 291 schoolbook volumes, which are comparable in content for the five idioms. We use automatic alignment methods to extract 207k multi-parallel segments from the books, with more than 2M tokens in total. A small-scale human evaluation confirms that the segments are highly parallel, making the dataset suitable for NLP applications such as machine translation between Romansh idioms. We release the parallel and unaligned versions of the dataset under a CC-BY-NC-SA license and demonstrate its utility for machine translation by training and evaluating an LLM and a supervised multilingual MT model on the dataset.

小语种平行语料库机器翻译罗曼什语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。