arXiv:2602.04241cs.CL2026-02被引 1

改进分词方法可显著提升乌拉尔语系低资源语言的语法标注准确率。

Tokenization and Morphological Fidelity in Uralic NLP: A Cross-Lingual Evaluation

  • 采用重叠BPE分词法,减少词素碎片化,更好保留词形信息。
  • 在拉丁字母语言中,新方法使词性标注准确率明显提升。
  • 适合研究黏着语、低资源语言的跨语言迁移任务。

子词分词对自然语言处理性能有关键影响,但在形态丰富且资源匮乏的语言家族中仍研究不足。本研究系统比较了三种子词范式——字节对编码(BPE)、重叠BPE(OBPE)和单语模型(Unigram)——在六种乌拉尔语系语言上的表现,涵盖不同资源水平与类型多样性。以词性标注为控制下游任务,结果表明:在拉丁字母组中,OBPE在词形对齐度和标注准确率上均优于传统方法,主要得益于开放类词的碎片化减少及频次分布更均衡。迁移效果还受下游标注架构影响,与训练量和谱系亲缘性交互作用。综合来看,形态敏感的分词不仅是预处理选择,更是实现黏着语、低资源语言有效跨语言迁移的关键因素。

原文摘要 · Abstract (English)

Subword tokenization critically affects Natural Language Processing (NLP) performance, yet its behavior in morphologically rich and low-resource language families remains under-explored. This study systematically compares three subword paradigms -- Byte Pair Encoding (BPE), Overlap BPE (OBPE), and Unigram Language Model -- across six Uralic languages with varying resource availability and typological diversity. Using part-of-speech (POS) tagging as a controlled downstream task, we show that OBPE consistently achieves stronger morphological alignment and higher tagging accuracy than conventional methods, particularly within the Latin-script group. These gains arise from reduced fragmentation in open-class categories and a better balance across the frequency spectrum. Transfer efficacy further depends on the downstream tagging architecture, interacting with both training volume and genealogical proximity. Taken together, these findings highlight that morphology-sensitive tokenization is not merely a preprocessing choice but a decisive factor in enabling effective cross-lingual transfer for agglutinative, low-resource languages.

分词乌拉尔语低资源形态学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。