用因果干预揭示变压器模型中语法岛的渐变阻断机制
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs

- 通过因果干预定位模型中与语法相关的关键子空间
- 发现连词'and'在可提取与不可提取结构中表征不同
- 为语言学提供新假设,适合关注神经语言学的研究者
我们展示因果干预如何揭示英语句法中的长期难题——语法岛现象。从并列动词短语中提取信息常受抑制,但可接受性随词汇内容呈现梯度变化(如:'I know what he hates art and loves' vs. 'I know what he looked down and saw')。现代变压器语言模型复现了人类判断的这种梯度。通过隔离变压器块、注意力模块和MLP中功能相关的子空间,我们证明从并列岛中提取依赖于与标准wh-依赖相同的填充-间隙机制,但这些机制被不同程度地选择性阻断。通过对大规模无关文本投影到这些因果识别的子空间,我们提出一个新语言学假说:连词'and'在可提取与不可提取结构中表征不同,分别对应关系依赖表达与纯连接用法。结果表明,机械可解释性可推动句法研究,生成关于语言表征与处理的新假设。
原文摘要 · Abstract (English)
We show how causal interventions in Transformer models provide insights into English syntax by focusing on a long-standing challenge for syntactic theory: syntactic islands. Extraction from coordinated verb phrases is often degraded, yet acceptability varies gradiently with lexical content (e.g., "I know what he hates art and loves" vs. "I know what he looked down and saw"). We show that modern Transformer language models replicate human judgments across this gradient. Using causal interventions that isolate functionally relevant subspaces in Transformer blocks, attention modules, and MLPs, we demonstrate that extraction from coordination islands engages the same filler-gap mechanisms as canonical wh-dependencies, but that these mechanisms are selectively blocked to varying degrees. By projecting a large corpus of unrelated text onto these causally identified subspaces, we derive a novel linguistic hypothesis: the conjunction "and" is represented differently in extractable versus non-extractable constructions, corresponding to expressions encoding relational dependencies versus purely conjunctive uses. These results illustrate how mechanistic interpretability can inform syntax, generating new hypotheses about linguistic representation and processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。