arXiv:2601.04664cs.CLcs.AI2026-01

通过干预分析,精准定位多语言模型中真正影响特定语言的神经元。

CRANE: Causal Relevance Analysis of Language-Specific Neurons in Multilingual Large Language Models

  • 基于功能必要性而非激活强度,识别语言相关神经元。
  • 屏蔽目标语言神经元会显著降级该语言表现,其他语言基本不受影响。
  • 适用于研究多语言模型内部机制,尤其适合关注语言特异性的人。

多语言大模型在多种语言上表现优异,但其神经元层面的语言能力组织方式仍不明确。以往工作主要依赖激活强度判断语言相关神经元,易混淆语言偏好与功能重要性。本文提出CRANE,一种基于功能必要性的分析框架,通过针对性神经元干预重新定义语言特异性,以神经元对语言条件预测的贡献度衡量其专属性。实验在英文、中文和越南语多个基准上进行,结合专用相关性指标与基础模型到对话模型的迁移分析,结果表明CRANE比激活基方法更精确地分离出语言特异性组件。神经元干预揭示出一致的非对称模式:屏蔽某语言相关神经元会显著降低该语言性能,而其他语言性能保持相对稳定,说明神经元具有语言选择性但非排他性。代码将公开。

原文摘要 · Abstract (English)

Multilingual large language models (LLMs) achieve strong performance across languages, yet how language capabilities are organized at the neuron level remains poorly understood. Prior work has identified language-related neurons mainly through activation-based heuristics, which conflate language preference with functional importance. We propose CRANE, a relevance-based analysis framework that redefines language specificity in terms of functional necessity, identifying language-specific neurons through targeted neuron-level interventions. CRANE characterizes neuron specialization by their contribution to language-conditioned predictions rather than activation magnitude. Our implementation will be made publicly available. Neuron-level interventions reveal a consistent asymmetric pattern: masking neurons relevant to a target language selectively degrades performance on that language while preserving performance on other languages to a substantial extent, indicating language-selective but non-exclusive neuron specializations. Experiments on English, Chinese, and Vietnamese across multiple benchmarks, together with a dedicated relevance-based metric and base-to-chat model transfer analysis, show that CRANE isolates language-specific components more precisely than activation-based methods.

多语言模型神经元分析因果干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。