用大模型自动生成濒危语言语法,仅需双语词典和例句。
Can LLMs Help Create Grammar?: Automating Grammar Creation for Endangered Languages with In-Context Learning
- 通过上下文学习,仅用双语词典生成正式语法规则。
- 在莫克伦语上成功捕捉核心语法结构和词汇信息。
- 为资源匮乏语言提供低成本语言保存方案。
在当今濒危语言的记录与保护中,大型语言模型(LLMs)的应用展现出巨大潜力。本文探讨了大模型(尤其是上下文学习)如何协助生成低资源语言的语法信息,这些语言数据有限。以莫克伦语为例,评估了大模型在仅有双语词典和平行句子的情况下,无需从头训练模型即可生成连贯语法规则和词汇条目的有效性。方法包括组织现有语言数据并设计提示,以高效生成正式的XLE语法。结果表明,大模型能成功捕捉关键语法结构和词汇信息,但仍存在英语语法偏见等挑战。该研究突显了大模型在增强语言记录方面的潜力,为生成语言数据提供成本效益高的解决方案,助力濒危语言的保护。
原文摘要 · Abstract (English)
Yes! In the present-day documenting and preserving endangered languages, the application of Large Language Models (LLMs) presents a promising approach. This paper explores how LLMs, particularly through in-context learning, can assist in generating grammatical information for low-resource languages with limited amount of data. We takes Moklen as a case study to evaluate the efficacy of LLMs in producing coherent grammatical rules and lexical entries using only bilingual dictionaries and parallel sentences of the unknown language without building the model from scratch. Our methodology involves organising the existing linguistic data and prompting to efficiently enable to generate formal XLE grammar. Our results demonstrate that LLMs can successfully capture key grammatical structures and lexical information, although challenges such as the potential for English grammatical biases remain. This study highlights the potential of LLMs to enhance language documentation efforts, providing a cost-effective solution for generating linguistic data and contributing to the preservation of endangered languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。