arXiv:2505.15722cs.CLcs.AI2025-05EMNLP被引 4

发现相似语言间低资源语种更易记忆,揭示多语言模型记忆机制新规律。

Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities

  • 基于语言相似性构建图模型,分析跨语言记忆关联
  • 低资源相似语言反而记忆率更高,突破数据量决定论
  • 为多语言模型安全与迁移能力提供新视角,适合模型评估者

本文首次系统研究多语言大模型(MLLMs)中的记忆现象,覆盖95种语言,涵盖不同模型规模、架构及记忆定义。随着多语言模型广泛应用,理解其记忆行为至关重要。然而以往研究多聚焦单语模型,忽视了训练语料固有的长尾分布。我们发现,传统认为记忆程度与训练数据量强相关的假设无法完全解释多语言模型的记忆模式。我们提出一种基于图的跨语言相关性度量方法,引入语言相似性来分析跨语言记忆。分析显示,在语言相似的群体中,训练样本较少的语言反而表现出更高的记忆率,这一趋势仅在显式建模跨语言关系时显现。结果表明,语言感知视角对评估和缓解多语言模型的记忆漏洞至关重要。该研究还提供了实证证据:语言相似性既解释了多语言模型的记忆现象,也支撑了跨语言迁移能力,对多语言自然语言处理具有广泛意义。

原文摘要 · Abstract (English)

We present the first comprehensive study of Memorization in Multilingual Large Language Models (MLLMs), analyzing 95 languages using models across diverse model scales, architectures, and memorization definitions. As MLLMs are increasingly deployed, understanding their memorization behavior has become critical. Yet prior work has focused primarily on monolingual models, leaving multilingual memorization underexplored, despite the inherently long-tailed nature of training corpora. We find that the prevailing assumption, that memorization is highly correlated with training data availability, fails to fully explain memorization patterns in MLLMs. We hypothesize that the conventional focus on monolingual settings, effectively treating languages in isolation, may obscure the true patterns of memorization. To address this, we propose a novel graph-based correlation metric that incorporates language similarity to analyze cross-lingual memorization. Our analysis reveals that among similar languages, those with fewer training tokens tend to exhibit higher memorization, a trend that only emerges when cross-lingual relationships are explicitly modeled. These findings underscore the importance of a \textit{language-aware} perspective in evaluating and mitigating memorization vulnerabilities in MLLMs. This also constitutes empirical evidence that language similarity both explains Memorization in MLLMs and underpins Cross-lingual Transferability, with broad implications for multilingual NLP.

多语言模型记忆机制语言相似性模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。