arXiv:2410.21353cs.CLcs.AI2024-10被引 2

通过因果句分析GPT-2如何从语法推断语义

Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics

  • 用明确因果句测试模型内部机制,定位语法线索
  • 前2-3层处理因果语法,后期层关注语义敏感性
  • 为理解大模型推理提供可干预的分析路径

尽管可解释性研究揭示了部分Transformer类大模型的内部算法,但自然语言推理因深层上下文依赖和模糊性,难以被清晰分类。构建适合因果干预的域内与域外例子以提出明确问题仍具挑战。虽然已有大量研究聚焦特定任务(如间接宾语识别),但通过电路分析解码自然语言推理仍因复杂性困难重重。本文以‘我打开伞因为下雨了’等清晰因果句为例,利用GPT-2 small进行精心设计的干预实验,发现因果语法集中于前2-3层,而后期某些注意力头对因果句的荒谬变体表现出更高敏感度。这表明模型可能通过(1)检测句法线索,(2)激活末层特定注意力头来捕捉语义关系实现推理。

原文摘要 · Abstract (English)

While interpretability research has shed light on some internal algorithms utilized by transformer-based LLMs, reasoning in natural language, with its deep contextuality and ambiguity, defies easy categorization. As a result, formulating clear and motivating questions for circuit analysis that rely on well-defined in-domain and out-of-domain examples required for causal interventions is challenging. Although significant work has investigated circuits for specific tasks, such as indirect object identification (IOI), deciphering natural language reasoning through circuits remains difficult due to its inherent complexity. In this work, we take initial steps to characterize causal reasoning in LLMs by analyzing clear-cut cause-and-effect sentences like "I opened an umbrella because it started raining," where causal interventions may be possible through carefully crafted scenarios using GPT-2 small. Our findings indicate that causal syntax is localized within the first 2-3 layers, while certain heads in later layers exhibit heightened sensitivity to nonsensical variations of causal sentences. This suggests that models may infer reasoning by (1) detecting syntactic cues and (2) isolating distinct heads in the final layers that focus on semantic relationships.

大模型推理因果分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。