用上下文无关文法生成纳瓦特尔语句子,扩充稀缺语料库。
A First Context-Free Grammar Applied to Nawatl Corpora Augmentation
- 构建纳瓦特尔语的上下文无关文法,自动生成语法正确句子。
- 生成语料使FastText在语义任务上表现优于部分大模型。
- 为资源匮乏的原住民语言提供可扩展的语料增强方案。
本文提出一种适用于纳瓦特尔语(Nawatl)的上下文无关文法(CFG)。纳瓦特尔语属于$π$-语言类型,数字资源极贫乏,可用于机器学习的语料几乎不存在。本研究旨在通过生成大量语法正确的合成句子,显著扩充语料库。所生成的语料库名为$π$- extsc{yalli},可用于训练FastText等算法,并在句子级语义任务上进行评估。初步结果显示,使用该文法后,在某些任务上相较部分大语言模型取得可比性提升。但观察表明,若想实现更显著改进,还需构建更精准刻画纳瓦特尔语特性的文法。
原文摘要 · Abstract (English)
In this article we introduce a context-free grammar (CFG) for the Nawatl language. Nawatl (or Nahuatl) is an Amerindian language of the $π$-language type, i.e. a language with few digital resources, in which the corpora available for machine learning are virtually non-existent. The objective here is to generate a significant number of grammatically correct artificial sentences, in order to increase the corpora available for language model training. We want to show that a grammar enables us significantly to expand a corpus in Nawatl which we call $π$-\textsc{yalli}. The corpus, thus enriched, enables us to train algorithms such as FastText and to evaluate them on sentence-level semantic tasks. Preliminary results show that by using the grammar, comparative improvements are achieved over some LLMs. However, it is observed that to achieve more significant improvement, grammars that model the Nawatl language even more effectively are required.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。