arXiv:2508.02241cs.CL2025-08中稿 · IJCNLP-AACL 2025被引 5

发现多语言模型中可独立调控的文化神经元

Isolating Culture Neurons in Multilingual Large Language Models

  • 通过新方法定位文化特异性神经元,分离其与语言神经元
  • 文化神经元主要分布于高层,且在6种文化中各具独立性
  • 为模型公平性与文化对齐提供可编辑的神经机制

语言与文化密不可分,但多语言大模型如何编码文化仍不明确。本文基于识别语言特异性神经元的方法,定位并分离文化特异性神经元,厘清其与语言神经元的重叠与交互关系。为支持实验,构建了包含8520万词元、覆盖六种文化的MUREL数据集。定位与干预实验表明,不同文化在模型中由不同的神经元群体编码,主要集中于上层;这些文化神经元可基本独立于语言神经元或其他文化神经元进行调节。结果表明,多语言模型中的文化知识与倾向可被选择性隔离与修改,对公平性、包容性及对齐具有重要意义。代码与数据见https://github.com/namazifard/Culture_Neurons。

原文摘要 · Abstract (English)

Language and culture are deeply intertwined, yet it has been unclear how and where multilingual large language models encode culture. Here, we build on an established methodology for identifying language-specific neurons to localize and isolate culture-specific neurons, carefully disentangling their overlap and interaction with language-specific neurons. To facilitate our experiments, we introduce MUREL, a curated dataset of 85.2 million tokens spanning six different cultures. Our localization and intervention experiments show that LLMs encode different cultures in distinct neuron populations, predominantly in upper layers, and that these culture neurons can be modulated largely independently of language-specific neurons or those specific to other cultures. These findings suggest that cultural knowledge and propensities in multilingual language models can be selectively isolated and edited, with implications for fairness, inclusivity, and alignment. Code and data are available at https://github.com/namazifard/Culture_Neurons.

文化神经元多语言模型可编辑性公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。