arXiv:2604.22166cs.CL2026-04ACL

发现语言模型处理句法依赖有共享神经机制,但否定词用法没有。

Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models

论文配图:Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models
图 1 · 摘自论文原文
  • 通过激活拼接技术定位注意力头和MLP块的功能角色。
  • 早期到中期层存在共享的填充-空缺依赖处理机制,而否定词使用无统一机制。
  • 该机制可泛化到分布外数据,适合关注模型内部句法理解的研究者。

尽管语言模型展现出复杂的句法能力,但其内部机制与语言学中跨结构原则的一致性仍不明确。本研究通过细粒度因果可解释性方法,探究模型在不同句法结构间是否采用共享神经机制。聚焦填充-空缺依赖和否定极性词(NPI)许可问题,利用激活拼接识别特定注意力头与MLP模块的功能角色。结果表明:填充-空缺依赖在早期至中期层存在高度局部化且共享的机制;而NPI处理则无此类统一机制。此外,激活拼接识别出的机制具有跨分布泛化能力,而监督式可解释方法在狭窄语言分布上易过拟合。最后,通过操纵所识别组件,验证了其对可接受性判断基准性能的提升效果。

原文摘要 · Abstract (English)

While language models demonstrate sophisticated syntactic capabilities, the extent to which their internal mechanisms align with cross-constructional principles studied in linguistics remains poorly understood. This study investigates whether models employ shared neural mechanisms across different syntactic constructions by applying causal interpretability methods at a granular level. Focusing on filler-gap dependencies and negative polarity item (NPI) licensing, we utilize activation patching to identify the functional roles of specific attention heads and MLP blocks. Our results reveal a highly localized and shared mechanism for filler-gap dependencies located in the early to middle layers, whereas NPI processing exhibits no such unified mechanism. Furthermore, we find that these mechanisms identified by activation patching generalize to out-of-distribution, while distributed alignment search, a supervised interpretability method, is susceptible to overfitting on narrow linguistic distributions. Finally, we validate our findings by demonstrating that the manipulation of the identified components improves model performance on acceptability judgment benchmarks.

句法分析可解释性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。