arXiv:2604.05090cs.CLcs.LG2026-04ACL被引 2

多语言模型更依赖字形而非语言结构组织表征。

Multilingual Language Models Encode Script Over Linguistic Structure

  • 通过熵度量和稀疏自编码器分析模型激活单元。
  • 罗马化输入产生几乎不重叠的表征,与母语或英文无关。
  • 深层才逐渐显现语言类型特征,生成最敏感于表面形式不变的单元。

多语言语言模型将语系和书写系统差异显著的语言表征统一在共享参数空间中,但其内部组织机制仍不明确。本文通过语言激活概率熵(LAPE)度量分析不同模型家族和规模下的语言相关单元,并利用稀疏自编码器分解激活模式。结果表明,这些单元强烈依赖书写形式:罗马化导致近乎不重叠的表征,且与原文字形输入或英语均不匹配;而词序打乱对单元身份影响有限。探测实验显示,语言类型结构在深层才逐渐可被识别;因果干预表明,生成过程对表面形式扰动不变的单元最敏感,而非仅依赖类型对齐的单元。总体而言,多语言模型以表面形式为基础组织表征,语言抽象性渐进浮现,未坍缩为统一的中间语。

原文摘要 · Abstract (English)

Multilingual language models (LMs) organize representations for typologically and orthographically diverse languages into a shared parameter space, yet the nature of this internal organization remains elusive. In this work, we investigate which linguistic properties - abstract language identity or surface-form cues - shape multilingual representations. To do so, we analyze language-associated units across different model families and scales using the Language Activation Probability Entropy (LAPE) metric, and further decompose activations with Sparse Autoencoders. We find that these units are strongly conditioned on orthography: romanization induces near-disjoint representations that align with neither native-script inputs nor English, while word-order shuffling has limited effect on unit identity. Probing shows that typological structure becomes increasingly accessible in deeper layers, while causal interventions indicate that generation is most sensitive to units that are invariant to surface-form perturbations rather than to units identified by typological alignment alone. Overall, our results suggest that multilingual LMs organize representations around surface form, with linguistic abstraction emerging gradually without collapsing into a unified interlingua.

多语言模型表征学习字形语言类型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。