arXiv:2412.12143cs.CLcs.AI2024-12被引 1

用斯瓦希里语迁移学习,提升科摩罗方言的自然语言处理能力

Harnessing Transfer Learning from Swahili: Advancing Solutions for Comorian Dialects

  • 用斯瓦希里语数据通过词法距离筛选,构建混合科摩罗语数据集
  • 机器翻译模型ROUGE-1达0.6826,语音识别词错率39.50%
  • 为濒危语言提供可复用的NLP解决方案,适合低资源语言研究者

尽管斯瓦希里语等非洲语言已有充足资源开发高性能自然语言处理系统,但大陆其他许多语言仍缺乏此类支持。针对这些处于初期阶段的语言,迁移学习提供了可行路径——利用语义相近语言的良好表示来提升低资源语言性能。本文首次探索为科摩罗语(属班图语系的四种语言/方言)构建NLP技术。基于人类跨语言理解的启发,我们提出将斯瓦希里语数据与科摩罗语混合,并仅保留词法距离最近的数据片段。在自动语音识别(ASR)和机器翻译(MT)两个任务上验证该方法:MT模型取得ROUGE-1、ROUGE-2、ROUGE-L分别为0.6826、0.42、0.6532的成绩;ASR系统词错误率(WER)为39.50%,字符错误率(CER)为13.76%。该研究对推动欠代表语言的NLP发展具有重要意义,有助于在数字时代保护与推广科摩罗语语言遗产。

原文摘要 · Abstract (English)

If today some African languages like Swahili have enough resources to develop high-performing Natural Language Processing (NLP) systems, many other languages spoken on the continent are still lacking such support. For these languages, still in their infancy, several possibilities exist to address this critical lack of data. Among them is Transfer Learning, which allows low-resource languages to benefit from the good representation of other languages that are similar to them. In this work, we adopt a similar approach, aiming to pioneer NLP technologies for Comorian, a group of four languages or dialects belonging to the Bantu family. Our approach is initially motivated by the hypothesis that if a human can understand a different language from their native language with little or no effort, it would be entirely possible to model this process on a machine. To achieve this, we consider ways to construct Comorian datasets mixed with Swahili. One thing to note here is that in terms of Swahili data, we only focus on elements that are closest to Comorian by calculating lexical distances between candidate and source data. We empirically test this hypothesis in two use cases: Automatic Speech Recognition (ASR) and Machine Translation (MT). Our MT model achieved ROUGE-1, ROUGE-2, and ROUGE-L scores of 0.6826, 0.42, and 0.6532, respectively, while our ASR system recorded a WER of 39.50\% and a CER of 13.76\%. This research is crucial for advancing NLP in underrepresented languages, with potential to preserve and promote Comorian linguistic heritage in the digital age.

低资源语言迁移学习语音识别机器翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。