arXiv:2412.15821cs.CLcs.AI2024-12被引 1

为濒危语言纳瓦特尔语构建首个机器学习可用语料库

$π$-yalli: un nouveau corpus pour le nahuatl

  • 构建适配机器学习的纳瓦特尔语语料库π-YALLI
  • 支持开发词形统一、分词、词性标注等基础NLP工具
  • 适合语言保护与小语种自然语言处理研究者

NAHU²项目是法墨合作计划,旨在构建适配机器学习的π-YALLI语料库,以推动纳瓦特尔语计算资源的发展。纳瓦特尔语虽有约200万人使用,但缺乏计算资源。π-YALLI语料库将支持开展相关研究,用于开发语言模型(无论动态与否),进而促进自然语言处理工具的建设,包括:a) 字素统一器,b) 词切分器,c) 词性标注器,d) 基于内容的自动文本摘要;可能还包括e) 翻译器(基于概率或学习的方法)。

原文摘要 · Abstract (English)

The NAHU$^2$ project is a Franco-Mexican collaboration aimed at building the $π$-YALLI corpus adapted to machine learning, which will subsequently be used to develop computer resources for the Nahuatl language. Nahuatl is a language with few computational resources, even though it is a living language spoken by around 2 million people. We have decided to build $π$-YALLI, a corpus that will enable to carry out research on Nahuatl in order to develop Language Models (LM), whether dynamic or not, which will make it possible to in turn enable the development of Natural Language Processing (NLP) tools such as: a) a grapheme unifier, b) a word segmenter, c) a POS grammatical analyser, d) a content-based Automatic Text Summarization; and possibly, e) a translator translator (probabilistic or learning-based).

小语种语料库语言保护NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。