对濒危语言纳瓦特尔语,可控数据复制能提升语义嵌入效果。
Corpora deduplication or duplication in Natural Language Processing of few resourced languages ? A case of study: The Mexico's Nahuatl

- 通过增量复制方式扩展小语种语料库,提升模型训练质量。
- 在句子级语义相似度任务中,性能有适度提升。
- 首次应用于纳瓦特尔语,适合资源匮乏语言研究者参考。
本文探讨在计算资源有限的语言(即$π$-语言)中,数据复制是否有助于自然语言处理。以讲者超200万人、方言众多的纳瓦特尔语为例,其可用于训练大语言模型的语料几乎不存在。研究目标是通过可控方式扩展现有$π$-yalli语料库,采用增量复制技术训练静态词向量,并在句子级语义相似度任务中评估性能。结果表明,与原始语料相比,增量复制带来适度性能提升。据我们所知,该方法尚未在相关文献中应用。
原文摘要 · Abstract (English)
In this article, we seek to answer the following question: could data duplication be useful in Natural Language Processing (NLP) for languages with limited computational resources? In this type of languages (or $π$-languages), corpora available for training Large Language Models are virtually non-existent. In particular, we will study the impact of corpora expansion in Nawatl, an agglutinative and polysynthetic $π$-language spoken by over 2 million people, with a large number of dialectal varieties. The aim is to expand the new $π$-yalli corpus, which contains a limited number of Nawatl texts, by duplicating it in a controlled way. In our experiments, we will use the incremental duplication technique. The aim is to learn embeddings that are well-suited to NLP tasks. Thus, static embeddings were trained and evaluated in a sentence-level semantic similarity task. Our results show a moderate improvement in performance when using incremental duplication compared to the results obtained using only the corpus without expansion. Furthermore, to our knowledge, this technique has not yet been used in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。