构建多语言数据集CUTE,提升低资源语言的跨语言迁移能力
CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages
- 用机器翻译构建中、英、维、藏四语语料库,含平行与非平行数据
- 数据集达50GB,是目前维吾尔语和藏语最大的开源语料库
- 验证了机器翻译质量接近中英水平,适合低资源语言研究
大型语言模型在多种自然语言处理任务中表现出卓越的零样本能力,显著提升用户体验与效率。然而,这一优势主要局限于资源丰富语言。对于众多低资源语言,支持仍显不足,训练语料匮乏被认为是主要原因。本文构建并开源了CUTE中文、维吾尔语、藏语、英语数据集,包含两套各25GB的四语言语料(一套平行、一套非平行),均通过机器翻译获得。CUTE涵盖两种资源丰富语言(中文、英文)和两种低资源语言(维吾尔语、藏语)。在构建前,人工评估表明中文-维吾尔语及中文-藏语的机器翻译质量已接近中文-英文水平。CUTE是迄今最大规模的维吾尔语与藏语开源语料库,实证其能有效增强大模型对低资源语言的处理能力,并探究了语料平行性在跨语言迁移学习中的作用。相关语料与模型已公开共享。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate exceptional zero-shot capabilities in various NLP tasks, significantly enhancing user experience and efficiency. However, this advantage is primarily limited to resource-rich languages. For the diverse array of low-resource languages, support remains inadequate, with the scarcity of training corpora considered the primary cause. We construct and open-source CUTE Chinese, Uyghur, Tibetan,English dataset, consisting of two 25GB sets of four-language corpora (one parallel and one non-parallel), obtained through machine translation. CUTE encompasses two resource-rich languages (Chinese and English) and two low-resource languages (Uyghur and Tibetan). Prior to constructing CUTE, human assessment validates that the machine translation quality between Chinese-Uyghur and Chinese-Tibetan approaches that of Chinese-English translation. CUTE represents the largest open-source corpus for Uyghur and Tibetan languages to date, and we demonstrate its effectiveness in enhancing LLMs' ability to process low-resource languages while investigating the role of corpus parallelism in cross-lingual transfer learning. The CUTE corpus and related models are made publicly available to the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。