arXiv:2607.04895cs.CL2026-07
构建首个符合通用依存标注标准的奥塞梯语词法语料库并训练高精度分析器。
Ossetic-COT: Designing a morphologically annotated corpus and morphological analyzer for Ossetic
- 基于语料库设计并标注奥塞梯语词法结构,符合通用依存框架。
- 训练的BERT模型在词性与形态标签上准确率达95.60%。
- 为小语种自然语言处理提供可用资源,适合语言学与低资源NLP研究者。
本文首次构建了符合通用依存标注标准的伊罗尼奥塞梯语形态标注语料库。该语料库源自《口头文本奥塞梯语语料库》,包含5454条人工标注句子,共74032个词项。我们利用此语料库训练了一个基于BERT的词法分析器,其在词性与形态标签上的准确率达到95.60%。该成果为奥塞梯语的自然语言处理提供了首个高质量形态分析工具,推动小语种语言技术发展。
原文摘要 · Abstract (English)
In this work we present the first morphologically annotated corpus for Iron Ossetic that conforms to the Universal Dependencies schema. The corpus includes 5454 manually annotated sentences from the Iron Ossetic Corpus of Oral Texts, containing 74032 tokens. We use this corpus to train a BERT-based morphological analyzer. The analyzer achieves tag accuracy of 95.60%.
小语种词法分析依存标注
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。