多语言翻译模型的注意力机制被无关符号主导,导致分析结果失真。
Attention Sinks in Massively Multilingual Neural Machine Translation:Discovery, Analysis, and Mitigation

- 发现非内容符号占注意力总量83%~91%,形成‘注意力黑洞’
- 原始数据低估内容相似性近一半,过滤后相似度从36.7%升至70.7%
- 提出过滤工具包,适用于非洲语言与非非洲语言的可靠可解释性研究
在NLLB-200(600M)的跨注意力分析中,我们发现一种系统性偏差:非内容标记——主要是句末符号、语言标签和标点——占据了总注意力质量的83%至91%。我们称其为“注意力黑洞”,这一现象扩展了对大语言模型的研究成果,并揭示其根源在于词汇设计而非位置偏见。该偏差导致原始指标将内容级相似性低估近一半(36.7%原始值对比70.7%过滤后值),使未经校正的分析不可靠。为此,我们验证了一种仅保留内容词的过滤方法,移除非内容标记并重新归一化分布。在1,000组平行语料(包括斯瓦希里语、基库尤语、索马里语、卢奥语等非洲语言及德语、土耳其语、中文、印地语等非非洲语言)上应用该方法,确认该现象具有普遍性,并恢复了被掩盖的语言学信号:教师强制模式与生成模式间存在16.9个百分点的差异,注意力熵显示语言家族聚类清晰,还揭示了索马里语主宾谓语序与单调对齐之间的隐藏关联。我们发布了过滤工具包与修正数据集,以支持多语言NMT可复现的可解释性研究。
原文摘要 · Abstract (English)
Cross-attention patterns in neural machine translation (NMT) are widely used to study how multilingual models align linguistic structure. We report a systematic artifact in cross-attention analysis of NLLB-200 (600M): non-content tokens - primarily end-of-sequence tokens, language tags, and punctuation - capture 83 percent to 91 percent of total cross-attention mass. We term these "attention sinks," extending findings from LLMs [Xiao et al., 2023] to NMT cross-attention and identifying a causal mechanism rooted in vocabulary design rather than position bias. This artifact causes raw metrics to underestimate content-level similarity by nearly half (36.7 percent raw vs. 70.7 percent filtered), rendering uncorrected analyses unreliable. To address this, we validate a content-only filtering methodology that removes non-content tokens and renormalizes the distribution. Applying this to 1,000 parallel sentences across African languages (Swahili, Kikuyu, Somali, Luo) and non-African benchmarks (German, Turkish, Chinese, Hindi), we confirm the artifact is universal and recover masked linguistic signals: a 16.9 percentage-point gap between teacher-forcing and generation modes, clear language-family clustering in attention entropy, and a hidden Somali paradox linking SOV word order to monotonic alignment. We release our filtering toolkit and corrected datasets to support reproducible interpretability research on multilingual NMT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。