arXiv:2505.10081cs.CL2025-05

首次系统探查非洲语言模型的内部知识,揭示其语言信息编码机制。

Designing and Contextualising Probes for African Languages

  • 针对六种非洲语言训练分层探测器,分析语言特征分布。
  • 发现适配非洲语言的模型比多语言模型编码更多语言信息。
  • 控制任务验证结果反映模型内在知识,非探针记忆。

针对非洲语言的预训练语言模型(PLMs)持续改进,但其进步原因仍不明确。本文首次系统研究了非洲语言PLMs的语言知识探查。我们为六种语言类型多样的非洲语言训练分层探测器,分析语言特征的分布情况,并为MasakhaPOS数据集设计控制任务以解释探针性能。结果表明,针对非洲语言微调的模型比大规模多语言模型更能编码目标语言的语义信息。我们的发现再次确认:词级句法信息集中在中后层,而句级语义信息则分布在所有层。通过控制任务和基线探查,我们证实探针性能反映模型内部知识,而非探针记忆。本研究将成熟可解释性技术应用于非洲语言模型,揭示了主动学习与多语言适应等策略成功背后的内在机制。

原文摘要 · Abstract (English)

Pretrained language models (PLMs) for African languages are continually improving, but the reasons behind these advances remain unclear. This paper presents the first systematic investigation into probing PLMs for linguistic knowledge about African languages. We train layer-wise probes for six typologically diverse African languages to analyse how linguistic features are distributed. We also design control tasks, a way to interpret probe performance, for the MasakhaPOS dataset. We find PLMs adapted for African languages to encode more linguistic information about target languages than massively multilingual PLMs. Our results reaffirm previous findings that token-level syntactic information concentrates in middle-to-last layers, while sentence-level semantic information is distributed across all layers. Through control tasks and probing baselines, we confirm that performance reflects the internal knowledge of PLMs rather than probe memorisation. Our study applies established interpretability techniques to African-language PLMs. In doing so, we highlight the internal mechanisms underlying the success of strategies like active learning and multilingual adaptation.

语言模型非洲语言可解释性探针分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。