用两种新语法生成大量合法纳瓦特尔语句子,扩充低资源语言数据集。
Two CFG Nahuatl for automatic corpora expansion
- 构建两种纳瓦特尔语上下文无关语法,用于生成合法句子。
- 人工扩展后语料使词向量在语义相似度任务中性能提升。
- 低成本词向量表现优于部分大模型,适合低资源语言研究。
本文提出两种用于纳瓦特尔语语料库扩展的上下文无关语法(CFG)。纳瓦特尔语是墨西哥的原住民语言,属于π-语言类型,数字化资源极少,现有用于大语言模型(LLMs)训练的语料几乎不存在。为此,我们引入两种新的纳瓦特尔语CFG,并以生成模式使用,旨在产生大量句法正确的虚构纳瓦特尔语句子,从而扩充语料库,用于学习非上下文嵌入。实验表明,相比仅使用原始语料的模型,经过人工扩展后的语料显著提升了嵌入表示的性能;同时,经济高效的词向量模型在语义相似度任务中的表现甚至优于某些大语言模型。
原文摘要 · Abstract (English)
The aim of this article is to introduce two Context-Free Grammars (CFG) for Nawatl Corpora expansion. Nawatl is an Amerindian language (it is a National Language of Mexico) of the $π$-language type, i.e. a language with few digital resources. For this reason the corpora available for the learning of Large Language Models (LLMs) are virtually non-existent, posing a significant challenge. The goal is to produce a substantial number of syntactically valid artificial Nawatl sentences and thereby to expand the corpora for the purpose of learning non contextual embeddings. For this objective, we introduce two new Nawatl CFGs and use them in generative mode. Using these grammars, it is possible to expand Nawatl corpus significantly and subsequently to use it to learn embeddings and to evaluate their relevance in a sentences semantic similarity task. The results show an improvement compared to the results obtained using only the original corpus without artificial expansion, and also demonstrate that economic embeddings often perform better than some LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。