用零样本检测任务分析多语言模型如何区分语言形式与语义。
Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
- 设计无训练的最小差异判别任务,探测模型对语言特征的敏感度。
- 语言区分能力随训练减弱并集中于浅层,语义区分则增强并稳定于深层。
- 适合研究模型内部表征结构或评估多语言学习进展的人参考。
我们提出一系列无需训练的ABX风格判别任务,用于评估多语言语言模型对语言身份(形式)和语义内容(意义)的表征。受语音处理启发,这些零样本任务衡量表示中微小差异是否可被可靠检测,为探针方法提供灵活且可解释的替代方案。在XLM-R(Conneau等,2020)不同预训练检查点与层间应用后发现,语言区分能力随训练下降并集中于低层,而语义区分能力持续增强并在深层稳定。进一步通过探针任务验证,发现我们的度量与语言学习表现存在部分一致性。结果表明,ABX任务构成了一种轻量级框架,可用于分析多语言表征的结构。
原文摘要 · Abstract (English)
We introduce a set of training-free ABX-style discrimination tasks to evaluate how multilingual language models represent language identity (form) and semantic content (meaning). Inspired from speech processing, these zero-shot tasks measure whether minimal differences in representation can be reliably detected. This offers a flexible and interpretable alternative to probing. Applied to XLM-R (Conneau et al, 2020) across pretraining checkpoints and layers, we find that language discrimination declines over training and becomes concentrated in lower layers, while meaning discrimination strengthens over time and stabilizes in deeper layers. We then explore probing tasks, showing some alignment between our metrics and linguistic learning performance. Our results position ABX tasks as a lightweight framework for analyzing the structure of multilingual representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。