揭秘双语模型如何从分语言学习到逐步融合
How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders
- 用稀疏自编码器追踪模型内部表征演化过程
- 中层位置出现明显的双语对齐,大模型更显著
- 可将训练后模型的分解表征注入中途模型提升性能
本研究探讨双语语言模型如何发展出复杂的内部表征。我们采用稀疏自编码器分析双语语言模型的内部表示,重点关注训练步数、层位置和模型规模的影响。分析表明,语言模型最初分别学习两种语言,随后在中层逐渐形成双语对齐。此外,发现这种双语倾向在更大模型中更明显。基于这些发现,我们提出一种新方法,将完全训练模型的分解表征注入中等训练阶段的模型,验证了双语表征对模型性能的关键作用。研究为理解语言模型如何获得双语能力提供了深入洞见。
原文摘要 · Abstract (English)
This study explores how bilingual language models develop complex internal representations. We employ sparse autoencoders to analyze internal representations of bilingual language models with a focus on the effects of training steps, layers, and model sizes. Our analysis shows that language models first learn languages separately, and then gradually form bilingual alignments, particularly in the mid layers. We also found that this bilingual tendency is stronger in larger models. Building on these findings, we demonstrate the critical role of bilingual representations in model performance by employing a novel method that integrates decomposed representations from a fully trained model into a mid-training model. Our results provide insights into how language models acquire bilingual capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。