arXiv:2606.24825cs.CLcs.LG2026-06

首个高质量马拉地语词性标注数据集,助力低资源语言NLP研究

L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models

论文配图:L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models
图 1 · 摘自论文原文
  • 构建32,354句马拉地语新闻文本的手动标注数据集
  • 最佳模型达88.67%词级准确率,15类标签宏平均F1为81.67%
  • 专为马拉地语设计,适合低资源语言与形态丰富语言研究者

词性标注是机器翻译、信息抽取和句法分析等自然语言处理任务的基础。尽管马拉地语有超过8300万人使用,位居全球前二十大语言之列,但其在标注语料库和标准化评估基准方面仍严重匮乏。马拉地语因丰富的形态变化、相对自由的词序、缺乏大小写规范以及与印地语和英语的普遍混用而对计算建模带来独特挑战。本文提出L3Cube-MahaPOS,一个基于新闻文本的高质量马拉地语词性标注数据集,包含32,354条手动标注句子,由精通马拉地语的标注员按照16类通用依存标注体系完成。通过涵盖Unicode归一化、德纳加里语感知分词和噪声过滤的预处理流程,确保所有数据集划分中的标签一致性。我们在六种模型家族上进行了基准测试:HMM、CRF、BiLSTM、BiLSTM+CharCNN、MuRIL和马拉地语专用Transformer模型MahaBERT-v2。最佳系统在15个标签类别上达到88.67%的词级准确率和81.67%的宏平均F1值。我们公开发布该数据集、标注指南和训练好的模型检查点,以推动马拉地语自然语言处理研究。

原文摘要 · Abstract (English)

Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing. Despite Marathi being spoken by over 83 million people and ranking among the top twenty most spoken languages worldwide, it remains severely under-resourced in annotated corpora and standardised evaluation benchmarks. Marathi presents unique challenges for computational modelling owing to its rich morphology, relatively free word order, lack of capitalisation conventions, and pervasive code-mixing with Hindi and English. We introduce L3Cube-MahaPOS, a gold-standard POS tagging dataset for Marathi comprising 32,354 manually annotated sentences drawn from news text. Annotation was performed entirely manually by a team of Marathi-proficient annotators following a 16-tag Universal Dependencies-aligned scheme. A structured preprocessing pipeline covering Unicode normalisation, Devanagari-aware tokenisation, and noise filtering ensures label consistency across all splits. We benchmark the dataset across six model families spanning HMM, CRF, BiLSTM, BiLSTM+CharCNN, MuRIL, and the Marathi-specific transformer MahaBERT-v2. The best system achieves 88.67\% token-level accuracy and a macro-F1 of 81.67% over 15 evaluated tag classes. We release the dataset, annotation guidelines, and trained model checkpoints to foster further research in Marathi NLP.

词性标注低资源语言马拉地语BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。