arXiv:2604.22723cs.LGcs.CL2026-04

用跨语言迁移和无监督聚类发现濒危班图语的词法特征

Zero-Shot Morphological Discovery in Low-Resource Bantu Languages via Cross-Lingual Transfer and Unsupervised Clustering

  • 结合斯瓦希里语迁移与聚类,自动识别词类和形态模式
  • 在91个标注词形下发现2455个词的词类,准确率达97.3%
  • 适合语言资源匮乏的班图语研究者使用

我们提出一种方法,通过跨语言迁移学习与无监督聚类相结合,发现低资源班图语的词法特征。针对仅有91个标注词形的吉里亚马语(nyf),该流程为2,455个词分配词类,并识别出两个此前未记录的形态模式:一类2词类的a-前缀变体(元音合并,wa- → a-,一致性95.1%)和一个收缩的k'-前缀(一致性98.5%)。对444个已知吉里亚马动词词形的外部验证显示78.2%的词干还原准确率;将v3语料库扩展至19,624词(9,014个唯一词干)后,在所有主要词类上达到97.3%的分词率和86.7%的词干还原率。该方法融合了来自斯瓦希里语的迁移学习(约60%词汇重叠)与聚类,通过加权投票整合互补优势:迁移擅长同源词识别,聚类则发现语言特异性创新。代码与生成的词典均已公开,以支持低资源班图语的词法建档。

原文摘要 · Abstract (English)

We present a method for discovering morphological features in low-resource Bantu languages by combining cross-lingual transfer learning with unsupervised clustering. Applied to Giriama (nyf), a language with only 91 labeled paradigms, our pipeline discovers noun class assignments for 2,455 words and identifies two previously undocumented morphological patterns: an a- prefix variant for Class 2 (vowel coalescence - the merger of two adjacent vowels - of wa-, 95.1% consistency) and a contracted k'- prefix (98.5% consistency). External validation on 444 known Giriama verb paradigms confirms 78.2% lemmatization accuracy, while a v3 corpus expansion to 19,624 words (9,014 unique lemmas) achieves 97.3% segmentation and 86.7% lemmatization rates across all major word classes. Our ensemble of transfer learning from Swahili and unsupervised clustering, combined via weighted voting, exploits complementary strengths: transfer excels at cognate detection (leveraging ~60% vocabulary overlap) while clustering discovers language-specific innovations invisible to transfer. We release all code and discovered lexicons to support morphological documentation for low-resource Bantu languages.

词法分析低资源语言迁移学习聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。