首个乌兹别克语词性标注数据集,基于BERT实现91%准确率。
BBPOS: BERT-based Part-of-Speech Tagging for Uzbek
- 用乌兹别克语BERT模型微调做词性标注
- 达91%平均准确率,优于多语言BERT和规则系统
- 能捕捉词缀带来的词性变化,适合低资源语言研究
本文推进了低资源乌兹别克语的自然语言处理研究,首次在乌兹别克语上评估两种未测试过的单语BERT模型,并构建首个公开可用的乌兹别克语统一词性标注(UPOS)基准数据集。微调后的模型在词性标注任务中达到91%的平均准确率,优于基线多语言BERT及规则型标注器。值得注意的是,这些模型能够通过词缀捕捉中间词性变化,展现上下文敏感性,而现有规则系统不具备此能力。
原文摘要 · Abstract (English)
This paper advances NLP research for the low-resource Uzbek language by evaluating two previously untested monolingual Uzbek BERT models on the part-of-speech (POS) tagging task and introducing the first publicly available UPOS-tagged benchmark dataset for Uzbek. Our fine-tuned models achieve 91% average accuracy, outperforming the baseline multi-lingual BERT as well as the rule-based tagger. Notably, these models capture intermediate POS changes through affixes and demonstrate context sensitivity, unlike existing rule-based taggers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。