开源610万词的多语言翻译数据集,助力低资源语言机器翻译
SMOL: Professionally translated parallel data for 115 under-represented languages
- 构建两大子数据集:覆盖广泛词汇的SMOLSENT与主题多元的SMOLDOC
- 在125个语言对上实现显著chrF提升,部分语言首次拥有公开翻译数据
- 附带事实性评分与理由,为多数语言提供首个事实性评估资源
我们开源了SMOL(Set of Maximal Overall Leverage),一套用于解锁低资源语言机器翻译的训练数据。SMOL已翻译成124种(持续增长)低资源语言,涵盖125个语言对,其中许多语言此前无公开资源,总计包含610万翻译词元。该数据集包含两个子集:SMOLSENT,选择具有广泛唯一词元覆盖的句子;SMOLDOC,聚焦于广泛主题覆盖的文档级资源。结合已发布的GATITOS,形成段落、句子和词元层级的完整内容组合。实验表明,使用SMOL提示或微调大语言模型可显著提升chrF得分。除翻译外,我们还为SMOLDOC中所有文档提供了事实性评分与推理依据,为大多数语言首次建立事实性数据集。
原文摘要 · Abstract (English)
We open-source SMOL (Set of Maximal Overall Leverage), a suite of training data to unlock machine translation for low-resource languages. SMOL has been translated into 124 (and growing) under-resourced languages (125 language pairs), including many for which there exist no previous public resources, for a total of 6.1M translated tokens. SMOL comprises two sub-datasets, each carefully chosen for maximum impact given its size: SMOLSENT, a set of sentences chosen for broad unique token coverage, and SMOLDOC, a document-level resource focusing on a broad topic coverage. They join the already released GATITOS for a trifecta of paragraph, sentence, and token-level content. We demonstrate that using SMOL to prompt or fine-tune Large Language Models yields robust chrF improvements. In addition to translation, we provide factuality ratings and rationales for all documents in SMOLDOC, yielding the first factuality datasets for most of these languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。