揭示英语主导模型语言混淆的内在机制并提出精准干预方法
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
- 通过神经元级分析定位语言切换关键位置与故障层
- 修改少量关键神经元可显著降低混淆率,保持语言流畅性
- 适合关注多语言模型可解释性与安全性的研究者
语言混淆——大型语言模型在用户意图下生成非预期语言——仍是核心挑战,尤其在以英语为中心的模型中。本文首次开展语言混淆的机械可解释性研究,结合行为基准测试与神经元级分析。利用语言混淆基准(LCB),我们发现语言切换点(CPs)是该现象的核心。通过分层分析(TunedLens)和靶向神经元归因,揭示最终层的转换失败导致混淆。进一步表明,基于多语言微调模型对比分析识别出的关键神经元编辑,能显著缓解混淆,同时基本保持通用能力与语言流畅性。该方法在多数语言上达到与多语言对齐相当的混淆降低效果,且输出更清晰、质量更高。研究为大模型内部动态提供了新见解,凸显神经元级干预在构建稳健、可解释的多语言建模中的潜力。代码与数据已开源:https://github.com/ercong21/lang_confusion。
原文摘要 · Abstract (English)
Language confusion -- where large language models (LLMs) generate unintended languages against the user's need -- remains a critical challenge, especially for English-centric models. We present the first mechanistic interpretability (MI) study of language confusion, combining behavioral benchmarking with neuron-level analysis. Using the Language Confusion Benchmark (LCB), we show that confusion points (CPs) -- specific positions where language switches occur -- are central to this phenomenon. Through layer-wise analysis with TunedLens and targeted neuron attribution, we reveal that transition failures in the final layers drive confusion. We further demonstrate that editing a small set of critical neurons, identified via comparative analysis with a multilingual-tuned counterpart, substantially mitigates confusion while largely preserving general competence and fluency. Our approach matches multilingual alignment in confusion reduction for many languages and yields cleaner, higher-quality outputs. These findings provide new insights into the internal dynamics of LLMs and highlight neuron-level interventions as a promising direction for robust, interpretable multilingual language modeling. Code and data are available at: https://github.com/ercong21/lang_confusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。