arXiv:2503.17456cs.CLcs.AI2025-03中稿 · ed被引 15

发现特定语言神经元无法提升低资源语言的跨语言性能。

Language-specific Neurons Do Not Facilitate Cross-Lingual Transfer

  • 通过激活熵和阈值法定位语言特异性神经元
  • 在XNLI和XQuAD任务上未显著提升低资源语言表现
  • 对多语言大模型跨语言泛化提出新挑战

多语言大语言模型旨在实现多种语言的鲁棒自然语言理解,但在低资源语言上的性能显著下降。本文探究现有识别语言特异性神经元的技术是否可用于提升低资源语言的跨语言任务表现。通过覆盖语言激活概率熵、基于激活概率的阈值法等方法,结合Llama 3.1和Mistral Nemo模型,进行神经元级LoRA微调实验。结果表明,此类神经元特定干预无法在下游任务(XNLI、XQuAD)中带来跨语言性能提升。研究揭示了跨语言泛化的挑战,并为多语言大模型提供了关键洞察。

原文摘要 · Abstract (English)

Multilingual large language models (LLMs) aim towards robust natural language understanding across diverse languages, yet their performance significantly degrades on low-resource languages. This work explores whether existing techniques to identify language-specific neurons can be leveraged to enhance cross-lingual task performance of lowresource languages. We conduct detailed experiments covering existing language-specific neuron identification techniques (such as Language Activation Probability Entropy and activation probability-based thresholding) and neuron-specific LoRA fine-tuning with models like Llama 3.1 and Mistral Nemo. We find that such neuron-specific interventions are insufficient to yield cross-lingual improvements on downstream tasks (XNLI, XQuAD) in lowresource languages. This study highlights the challenges in achieving cross-lingual generalization and provides critical insights for multilingual LLMs.

多语言模型神经元分析跨语言迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。