arXiv:2509.19343cs.CLcs.AI2025-09被引 1

首次为那加语构建分词标注系统,准确率达85.7%

Part-of-speech tagging for Nagamese Language using CRF

  • 用条件随机场建模,基于16112词的标注语料
  • 整体标注准确率85.7%,精确率86%,F1值85%
  • 填补资源匮乏语言NLP研究空白,适合语言保护项目参考

本文研究了那加语(Nagamese)的词性标注任务,该语言又称那加皮钦语,是东北印度那加人与阿萨姆人之间贸易交流中发展起来的一种以阿萨姆语为词汇基础的克里奥尔语。尽管英语、印地语等资源丰富语言已有大量词性标注研究,但那加语尚未开展相关工作。据我们所知,这是首次针对那加语进行的词性标注尝试。研究目标是为给定句子中的每个词标注其词性。为此构建了一个包含16,112个词元的标注语料库,并采用条件随机场(CRF)这一机器学习方法。实验结果表明,使用CRF实现了85.70%的整体标注准确率,精确率为86%,召回率为85%,F1得分为85%。

原文摘要 · Abstract (English)

This paper investigates part-of-speech tagging, an important task in Natural Language Processing (NLP) for the Nagamese language. The Nagamese language, a.k.a. Naga Pidgin, is an Assamese-lexified Creole language developed primarily as a means of communication in trade between the Nagas and people from Assam in northeast India. A substantial amount of work in part-of-speech-tagging has been done for resource-rich languages like English, Hindi, etc. However, no work has been done in the Nagamese language. To the best of our knowledge, this is the first attempt at part-of-speech tagging for the Nagamese Language. The aim of this work is to identify the part-of-speech for a given sentence in the Nagamese language. An annotated corpus of 16,112 tokens is created and applied machine learning technique known as Conditional Random Fields (CRF). Using CRF, an overall tagging accuracy of 85.70%; precision, recall of 86%, and f1-score of 85% is achieved. Keywords. Nagamese, NLP, part-of-speech, machine learning, CRF.

词性标注那加语CRF低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。