用现代语言数据训练神经模型,成功还原了班图语系的古词结构。
Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data
- 用Transformer模型分析14种班图语的词形数据,提取词根嵌入
- 90.9%的顶级名词候选词与历史重建的原始班图词一致,12个动词词根也匹配
- 模型能识别出符合地理分类的词群,适合语言学和跨语言研究者
我们研究了仅基于现代形态数据训练的神经模型能否恢复跨语言词汇结构。利用BantuMorph v7(一种针对班图语词形模式的Transformer),分析14种东、南部班图语,提取名词和动词词根的编码嵌入,识别出728个名词和1,525个动词的跨语言同源词候选。与已有历史资源(BLR3:4,786个重建的原始班图词;ASJP基础词汇)对比,前11个名词候选中有10个(90.9%)与已重建的原始班图词一致,包括*-ntU 'person'(8种语言)、*gombe 'cow'(9种语言)、*mUn(9种语言)。动词方面,12个同源词与重建的原始班图词根吻合,如*-bon- 'see' 和 *-jIm- 'stand',均在广泛地理范围内出现。通过独立翻译模型(NLLB-600M)交叉验证,两模型均复现了与既定古特里分类一致的词群与系统发育关系(p < 0.01)。跨语言名词类分析显示,全部13个活跃词类在各语言间保持>0.83余弦相似度(类内 > 类间,p < 10^-9)。由于数据仅限于东、南部班图语,结果表明模型恢复的是与原始班图语一致的共享词汇结构,而非明确区分原始保留与后期区域创新。
原文摘要 · Abstract (English)
We investigate whether neural models trained exclusively on modern morphological data can recover cross-lingual lexical structure consistent with historical reconstruction. Using BantuMorph v7, a transformer over Bantu morphological paradigms, we analyze 14 Eastern and Southern Bantu languages, extract encoder embeddings for their noun and verb lemmas, and identify 728 noun and 1,525 verb cognate candidates shared across 5+ languages. Evaluating these candidates against established historical resources-the Bantu Lexical Reconstructions database (BLR3; 4,786 reconstructed Proto-Bantu forms) and the ASJP basic vocabulary-we confirm 10 of the top 11 noun candidates (90.9%) align with previously reconstructed Proto-Bantu forms, including *-ntU 'person' (8 languages), *gombe 'cow' (9 languages), and *mUn (9 languages). Extending to verbs, 12 verb cognates align with reconstructed Proto-Bantu roots, including *-bon- 'see' and *-jIm- 'stand', each attested across wide geographic ranges. Cross-model validation using an independent translation model (NLLB-600M) confirms these patterns: both models recover cognate clusters and phylogenetic groupings consistent with established Guthrie-zone classifications (p < 0.01). Cross-lingual noun class analysis reveals that all 13 productive classes maintain >0.83 cosine similarity across languages (within-class > between-class, p < 10^-9). Our dataset is restricted to Eastern and Southern Bantu, so we interpret these results as recovering shared Bantu lexical structure consistent with Proto-Bantu rather than definitively distinguishing Proto-Bantu retentions from later regional innovations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。