用1万字数据让低资源语言恢复声调符号,效果最佳的是字符级大模型。
Diacritic Restoration for Low-Resource Indigenous Languages: Case Study with Bribri and Cook Islands Māori
- 用字符级大模型分解字符字节,提升声调符号还原能力。
- 需约1万词数据才能达到可靠性能,少于该量级效果差。
- 适合关注原住民语言数字化的学者与社区项目组。
我们研究了对极低资源语言进行声调符号恢复的方法,这对自然语言处理至关重要。研究对象为哥斯达黎加的布里布里语(Bribri)和库克群岛毛利语(Cook Islands Māori)。本文:(i) 比较不同算法在低资源环境下的声调符号恢复表现,包括声调符号;(ii) 分析达到目标性能所需的最小数据量;(iii) 对比不同资源条件下的结果;(iv) 探索声调符号修正任务。结果表明,微调后的字符级大模型表现最佳,可能因其能将复杂字符分解为UTF-8字节表示。相比之下,大规模多语言模型在数据受限时表现较差。所有模型中,性能稳定出现在约10,000词数据量时。零样本方法在所有情况下均表现不佳。本研究既回应了语言社群的实际需求,也推动了低资源情境下模型性能与泛化能力的学术探讨。
原文摘要 · Abstract (English)
We present experiments on diacritic restoration, a form of text normalization essential for natural language processing (NLP) tasks. Our study focuses on two extremely under-resourced languages: Bribri, a Chibchan language spoken in Costa Rica, and Cook Islands Māori, a Polynesian language spoken in the Cook Islands. Specifically, this paper: (i) compares algorithms for diacritics restoration in under-resourced languages, including tonal diacritics, (ii) examines the amount of data required to achieve target performance levels, (iii) contrasts results across varying resource conditions, and (iv) explores the related task of diacritic correction. We find that fine-tuned, character-level LLMs perform best, likely due to their ability to decompose complex characters into their UTF-8 byte representations. In contrast, massively multilingual models perform less effectively given our data constraints. Across all models, reliable performance begins to emerge with data budgets of around 10,000 words. Zero-shot approaches perform poorly in all cases. This study responds both to requests from the language communities and to broader NLP research questions concerning model performance and generalization in under-resourced contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。