arXiv:2409.15827cs.CL2024-09被引 13

用心理语言学方法找出大模型中负责语言能力的神经元

Unveiling Language Competence Neurons: A Psycholinguistic Approach to Model Interpretability

  • 设计英语心理语言实验,探测模型在三类语言任务中的神经表示
  • 发现GPT-2-XL在声音-性别、隐含因果上具备类人能力,但声音-形状任务表现差
  • 首次实现神经元级解释,揭示特定语言能力与专属神经元的对应关系

随着大语言模型(LLMs)语言能力不断提升,理解其如何表征语言能力仍面临重大挑战。本研究采用适用于探查语言认知深层机制的心理语言学范式,针对英文三种任务——音形关联、音性关联和隐含因果——探究基于Transformer的模型在神经元层面的表征。结果表明,尽管GPT-2-XL在音形关联任务中表现不佳,但在音性关联与隐含因果任务中展现出类人能力。通过针对性神经元消融与激活操控,发现当模型具备某种语言能力时,存在对应的特化神经元;反之则无。该研究首次将心理语言学实验引入模型语言能力的神经元层级分析,为模型可解释性提供新粒度,并揭示了变压器架构语言能力的内部驱动机制。

原文摘要 · Abstract (English)

As large language models (LLMs) advance in their linguistic capacity, understanding how they capture aspects of language competence remains a significant challenge. This study therefore employs psycholinguistic paradigms in English, which are well-suited for probing deeper cognitive aspects of language processing, to explore neuron-level representations in language model across three tasks: sound-shape association, sound-gender association, and implicit causality. Our findings indicate that while GPT-2-XL struggles with the sound-shape task, it demonstrates human-like abilities in both sound-gender association and implicit causality. Targeted neuron ablation and activation manipulation reveal a crucial relationship: When GPT-2-XL displays a linguistic ability, specific neurons correspond to that competence; conversely, the absence of such an ability indicates a lack of specialized neurons. This study is the first to utilize psycholinguistic experiments to investigate deep language competence at the neuron level, providing a new level of granularity in model interpretability and insights into the internal mechanisms driving language ability in the transformer-based LLM.

模型可解释性神经元分析语言能力心理语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。