arXiv:2507.18918cs.CLcs.AI2025-07被引 1

发现大模型对低资源语言激活不足,通过稀疏编码器定位问题并优化。

Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders

  • 用稀疏自编码器分析26层中10种语言的激活模式
  • 低资源语言早期层激活低26.27%,深层仍差19.89%
  • 基于激活微调后,马拉雅拉姆语激活提升87.69%

多语言大模型虽具跨语言泛化能力,但中低资源语言在ARC-Challenge、MMLU和HellaSwag等基准上表现较差。我们分析Gemma-2-2B在全部26个残差层中10种语言(包括中文、俄语、西班牙语、意大利语及印尼语、加泰罗尼亚语、马拉地语、马拉雅拉姆语、印地语)的激活模式,以英语为参照。利用稀疏自编码器(SAEs),揭示系统性差异:中低资源语言在早期层激活最多低26.27%,深层仍保持19.89%差距。针对此,采用基于激活的低秩适配(LoRA)微调,使马拉雅拉姆语激活提升87.69%,印地语提升86.32%,同时英语性能维持约91%。微调后基准测试显示小幅但稳定的提升,表明激活对齐是提升多语言模型性能的关键因素。

原文摘要 · Abstract (English)

Multilingual large language models (LLMs) exhibit strong cross-linguistic generalization, yet medium to low resource languages underperform on common benchmarks such as ARC-Challenge, MMLU, and HellaSwag. We analyze activation patterns in Gemma-2-2B across all 26 residual layers and 10 languages: Chinese (zh), Russian (ru), Spanish (es), Italian (it), medium to low resource languages including Indonesian (id), Catalan (ca), Marathi (mr), Malayalam (ml), and Hindi (hi), with English (en) as the reference. Using Sparse Autoencoders (SAEs), we reveal systematic disparities in activation patterns. Medium to low resource languages receive up to 26.27 percent lower activations in early layers, with a persistent gap of 19.89 percent in deeper layers. To address this, we apply activation-aware fine-tuning via Low-Rank Adaptation (LoRA), leading to substantial activation gains, such as 87.69 percent for Malayalam and 86.32 percent for Hindi, while maintaining English retention at approximately 91 percent. After fine-tuning, benchmark results show modest but consistent improvements, highlighting activation alignment as a key factor in enhancing multilingual LLM performance.

大模型多语言激活分析微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。