用随机森林提升纳瓦霍语等濒危语言的识别准确率
Is It Navajo? Accurate Language Detection in Endangered Athabaskan Languages
- 用随机森林模型识别纳瓦霍语及被误判的语言
- 对纳瓦霍语识别准确率达97%-100%
- 可扩展至其他阿塔巴斯克语族语言
濒危语言如纳瓦霍语(北美最广泛使用的原住民语言)在现代语言技术中严重缺失,加剧了其保护与复兴的困难。本研究评估了谷歌语言识别工具LangID,该工具目前不支持任何美洲原住民语言。为此,我们构建了一个基于纳瓦霍语及被LangID错误识别的二十种语言的随机森林分类器。尽管方法简单,但分类器在纳瓦霍语识别上达到了97%-100%的近乎完美准确率。此外,该模型在其他阿塔巴斯克语族语言中也表现出稳健性,表明其具有更广泛应用潜力。研究强调,需优先考虑语言多样性和适应性的NLP系统,而非集中化的通用方案,尤其在多元文化世界中支持弱势语言。本工作直接助力消除语言模型中的文化偏见,倡导开发服务于多样化语言社区的文化本地化NLP工具。
原文摘要 · Abstract (English)
Endangered languages, such as Navajo - the most widely spoken Native American language - are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. This study evaluates Google's Language Identification (LangID) tool, which does not currently support any Native American languages. To address this, we introduce a random forest classifier trained on Navajo and twenty erroneously suggested languages by LangID. Despite its simplicity, the classifier achieves near-perfect accuracy (97-100%). Additionally, the model demonstrates robustness across other Athabaskan languages - a family of Native American languages spoken primarily in Alaska, the Pacific Northwest, and parts of the Southwestern United States - suggesting its potential for broader application. Our findings underscore the pressing need for NLP systems that prioritize linguistic diversity and adaptability over centralized, one-size-fits-all solutions, especially in supporting underrepresented languages in a multicultural world. This work directly contributes to ongoing efforts to address cultural biases in language models and advocates for the development of culturally localized NLP tools that serve diverse linguistic communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。