为伊博语设计字形恢复框架,解决低资源语言文本歧义问题
Corpus-Based Approaches to Igbo Diacritic Restoration
- 基于n-gram、分类与嵌入三种模型构建字形恢复方法
- 利用上下文词序列预测缺失声调符号的正确形式
- 适用于低资源语言文本修复,尤其适合非洲本土语言
自然语言处理旨在让计算机识别并理解人类语言中的模式,但这一过程因语法、语用和语音等动态特性而复杂。当前研究多集中于英语、日语、德语、法语、俄语、中文等高资源语言,全球7000种语言中超过95%属于低资源语言,缺乏数据、工具与技术。本文综述了字形歧义问题及以往语言的消歧方法,聚焦伊博语,提出一个灵活的数据集生成框架。采用三种主要方法:标准n-gram模型通过前序词序列预测目标词的正确变体;分类模型使用目标词两侧词窗作为输入;嵌入模型则比较上下文词嵌入与候选变体向量的相似度得分。
原文摘要 · Abstract (English)
With natural language processing (NLP), researchers aim to enable computers to identify and understand patterns in human languages. This is often difficult because a language embeds many dynamic and varied properties in its syntax, pragmatics and phonology, which need to be captured and processed. The capacity of computers to process natural languages is increasing because NLP researchers are pushing its boundaries. But these research works focus more on well-resourced languages such as English, Japanese, German, French, Russian, Mandarin Chinese, etc. Over 95% of the world's 7000 languages are low-resourced for NLP, i.e. they have little or no data, tools, and techniques for NLP work. In this thesis, we present an overview of diacritic ambiguity and a review of previous diacritic disambiguation approaches on other languages. Focusing on the Igbo language, we report the steps taken to develop a flexible framework for generating datasets for diacritic restoration. Three main approaches, the standard n-gram model, the classification models and the embedding models were proposed. The standard n-gram models use a sequence of previous words to the target stripped word as key predictors of the correct variants. For the classification models, a window of words on both sides of the target stripped word was used. The embedding models compare the similarity scores of the combined context word embeddings and the embeddings of each of the candidate variant vectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。