arXiv:2410.13396cs.CL2024-10NAACL被引 1

用沙普利头值分析模型如何处理语法现象,发现相关注意力头会聚类。

Linguistically Grounded Analysis of Language Models using Shapley Head Values

  • 通过沙普利头值探测模型对语法结构的处理机制。
  • 发现处理同类语法现象的注意力头在模型中自然聚成簇。
  • 适合关注模型可解释性与语言理论对应关系的研究者。

理解语言模型中编码的语言知识对于提升其泛化能力至关重要。本文利用最近提出的基于沙普利头值(SHVs)的探测方法,研究模型对形态句法现象的处理。基于英语的BLiMP数据集,我们在BERT和RoBERTa两个主流模型上进行测试,比较它们对回指一致性和填充-空缺依赖等语言结构的处理方式。通过定量剪枝和定性聚类分析,我们发现负责处理相关语言现象的注意力头具有聚集特征。结果表明,基于SHV的归因揭示了两模型间不同的模式,为语言模型如何组织和处理语言信息提供了新见解。这些发现支持了语言模型学习到与语言理论对应的子网络的假设,对跨语言模型分析和自然语言处理中的可解释性具有潜在意义。

原文摘要 · Abstract (English)

Understanding how linguistic knowledge is encoded in language models is crucial for improving their generalisation capabilities. In this paper, we investigate the processing of morphosyntactic phenomena, by leveraging a recently proposed method for probing language models via Shapley Head Values (SHVs). Using the English language BLiMP dataset, we test our approach on two widely used models, BERT and RoBERTa, and compare how linguistic constructions such as anaphor agreement and filler-gap dependencies are handled. Through quantitative pruning and qualitative clustering analysis, we demonstrate that attention heads responsible for processing related linguistic phenomena cluster together. Our results show that SHV-based attributions reveal distinct patterns across both models, providing insights into how language models organize and process linguistic information. These findings support the hypothesis that language models learn subnetworks corresponding to linguistic theory, with potential implications for cross-linguistic model analysis and interpretability in Natural Language Processing (NLP).

可解释性注意力机制语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。