arXiv:2601.02303cs.CL2026-01被引 1

用机器学习区分30种纳瓦特尔方言,助力濒危语言数字化

Classifying several dialectal Nawatl varieties

  • 基于机器学习与神经网络构建方言分类模型
  • 针对30种纳瓦特尔方言及多种拼写形式进行识别
  • 为濒危语言资源建设提供技术方案,适合语言保护研究者

墨西哥拥有众多土著语言,其中纳瓦特尔语是使用人数最多的,目前约有超过两百万人使用(主要分布在北美和中美洲)。尽管其文化历史可追溯至15世纪,但该语言的计算机资源仍十分匮乏。问题在方言层面尤为突出:目前已确认约30种方言变体,且书面形式存在多种拼写方式。本文研究利用机器学习与神经网络技术,解决纳瓦特尔方言的自动分类问题。

原文摘要 · Abstract (English)

Mexico is a country with a large number of indigenous languages, among which the most widely spoken is Nawatl, with more than two million people currently speaking it (mainly in North and Central America). Despite its rich cultural heritage, which dates back to the 15th century, Nawatl is a language with few computer resources. The problem is compounded when it comes to its dialectal varieties, with approximately 30 varieties recognised, not counting the different spellings in the written forms of the language. In this research work, we addressed the problem of classifying Nawatl varieties using Machine Learning and Neural Networks.

语言识别机器学习濒危语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。