arXiv:2609.04819cs.CL2026-09

对比多种多语言可解释性方法,发现方向性偏差导致结果不一致,推荐使用ILO度量。

A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures

论文配图:A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures
图 1 · 摘自论文原文
  • 对比四种跨语言表征度量方法,分析其在21个模型中的表现差异。
  • 仅ILO与跨语言迁移性能相关性高达0.90,且不受模型规模和家族影响。
  • 建议报告ILO结果并附加方向性诊断,避免误导性结论。

多语言语言模型会发展出共享的跨语言表征,多种可解释性方法声称能量化这种共享。这些方法大多独立发展,当它们产生分歧时,无法判断是模型本身特性还是测量误差所致。本文在来自五个模型家族的21个基础模型(参数量125M-14B)上,比较了四种共享度量方法(CKA、ANC、GMM每词主导性、ILO),并将其与五个下游任务上的跨语言迁移性能进行相关性分析。结果发现,不同度量方法对模型中跨语言共享的量化存在差异,且分歧可能源于表征空间的方向性偏差(anisotropy)——即表征倾向于聚集在嵌入空间的一个狭窄锥形区域内。只有ILO与跨语言迁移性能的相关性(斯皮尔曼等级相关系数ρ = 0.90)在控制模型规模、模型家族和任务差异后依然显著。因此,我们推荐将ILO作为主要的共享度量指标,并配合方向性诊断一起报告。

原文摘要 · Abstract (English)

Multilingual language models develop shared cross-lingual representations, and various interpretability methods claim to quantify this sharing. These methods have been developed largely in isolation, and when they disagree, it is unclear whether the disagreement reflects a property of the model or an artifact of the measurement. We compare four sharing metrics (CKA, ANC, GMM dominance per token, and ILO) across 21 base models from five families (125M-14B parameters) and correlate each with cross-lingual transfer on five downstream tasks. We find that the metrics differ in their quantification of cross-lingual sharing in these models and suggest that the disagreement traces to anisotropy, the tendency of representations to cluster in a narrow cone of the embedding space. Only ILO's correlation with cross-lingual transfer (Spearman's $\rho = 0.90$) survives controls for model size, family, and per-task variation. We therefore recommend ILO as the primary sharing metric, to be reported alongside anisotropy diagnostics.

可解释性多语言方向性偏差表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。