arXiv:2501.10107cs.CLcs.AI2025-01被引 12

首个乌兹别克语词性标注数据集,基于BERT实现91%准确率。

BBPOS: BERT-based Part-of-Speech Tagging for Uzbek

  • 用乌兹别克语BERT模型微调做词性标注
  • 达91%平均准确率,优于多语言BERT和规则系统
  • 能捕捉词缀带来的词性变化,适合低资源语言研究

本文推进了低资源乌兹别克语的自然语言处理研究,首次在乌兹别克语上评估两种未测试过的单语BERT模型,并构建首个公开可用的乌兹别克语统一词性标注(UPOS)基准数据集。微调后的模型在词性标注任务中达到91%的平均准确率,优于基线多语言BERT及规则型标注器。值得注意的是,这些模型能够通过词缀捕捉中间词性变化,展现上下文敏感性,而现有规则系统不具备此能力。

原文摘要 · Abstract (English)

This paper advances NLP research for the low-resource Uzbek language by evaluating two previously untested monolingual Uzbek BERT models on the part-of-speech (POS) tagging task and introducing the first publicly available UPOS-tagged benchmark dataset for Uzbek. Our fine-tuned models achieve 91% average accuracy, outperforming the baseline multi-lingual BERT as well as the rule-based tagger. Notably, these models capture intermediate POS changes through affixes and demonstrate context sensitivity, unlike existing rule-based taggers.

词性标注低资源语言BERT乌兹别克语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。