针对埃塞俄比亚沃莱塔与戈法语的双语文本识别,提升跨语言社交内容治理能力。
Bilingual Word Level Language Identification for Omotic Languages
- 结合BERT与LSTM模型,融合上下文与序列特征提升识别精度。
- 在测试集上取得0.72的F1分数,验证方法有效性。
- 为非洲本土语言的多语处理提供可复用的技术基础。
语言识别旨在确定给定文本所使用的语言。在多语言社区中,文本常包含多种语言,因此双语语言识别(BLID)需区分同一文本中的两种语言。本文针对埃塞俄比亚南部的沃莱塔语和戈法语开展研究,由于两语言词汇存在相似性与差异性,识别难度较高。为此,我们对比多种方法,最终采用基于BERT的预训练模型与LSTM结合的方法,在测试集上达到0.72的F1分数。该成果有助于应对社交媒体中的非目标语言内容问题,并为该领域的后续研究提供基础支持。
原文摘要 · Abstract (English)
Language identification is the task of determining the languages for a given text. In many real world scenarios, text may contain more than one language, particularly in multilingual communities. Bilingual Language Identification (BLID) is the task of identifying and distinguishing between two languages in a given text. This paper presents BLID for languages spoken in the southern part of Ethiopia, namely Wolaita and Gofa. The presence of words similarities and differences between the two languages makes the language identification task challenging. To overcome this challenge, we employed various experiments on various approaches. Then, the combination of the BERT based pretrained language model and LSTM approach performed better, with an F1 score of 0.72 on the test set. As a result, the work will be effective in tackling unwanted social media issues and providing a foundation for further research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。