研究多语言模型训练中共享概念空间的形成与质量,揭示其早期出现但语言依赖性强。
When Meanings Meet: Investigating the Emergence and Quality of Shared Concept Spaces during Multilingual Language Model Training
- 通过激活修补法追踪模型训练过程中的跨语言概念表示
- 共享空间早期形成且持续优化,但对不同语言的对齐程度不同
- 翻译质量提升实为语义选择行为变化,非真正翻译能力增强
使用高多语言覆盖的大规模语言模型(LLM)训练日益重要,尤其在单语资源稀缺时。近期研究发现,LLM在处理多语言输入时会利用共享概念空间,有助于泛化与跨语言迁移。然而,以往研究常缺乏因果分析方法、深入错误分析或仅关注最终模型,未揭示这些空间如何在训练中逐步形成。本研究通过激活修补这一因果可解释性方法,考察EuroLLM预训练过程中语言无关概念空间的发展。我们分离出跨语言概念表示,并将其注入翻译提示中,以检验其能否独立于语言一致地改变翻译结果。结果显示,共享概念空间在训练初期即出现并持续优化,但其对齐具有语言依赖性。进一步细粒度人工分析表明,部分看似翻译质量提升实为行为变化——如对多义词选择特定语义、将跨语言同形异义词转译而非复制——而非真正翻译能力的提高。研究为跨语言对齐的训练动态提供了新见解,并揭示了因果可解释性方法在多语言场景下的适用条件。
原文摘要 · Abstract (English)
Training Large Language Models (LLMs) with high multilingual coverage is becoming increasingly important -- especially when monolingual resources are scarce. Recent studies have found that LLMs process multilingual inputs in shared concept spaces, thought to support generalization and cross-lingual transfer. However, these prior studies often do not use causal methods, lack deeper error analysis or focus on the final model only, leaving open how these spaces emerge during training. We investigate the development of language-agnostic concept spaces during pretraining of EuroLLM through the causal interpretability method of activation patching. We isolate cross-lingual concept representations, then inject them into a translation prompt to investigate how consistently translations can be altered, independently of the language. We find that shared concept spaces emerge early} and continue to refine, but that alignment with them is language-dependent}. Furthermore, in contrast to prior work, our fine-grained manual analysis reveals that some apparent gains in translation quality reflect shifts in behavior -- like selecting senses for polysemous words or translating instead of copying cross-lingual homographs -- rather than improved translation ability. Our findings offer new insight into the training dynamics of cross-lingual alignment and the conditions under which causal interpretability methods offer meaningful insights in multilingual contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。