厘清'机制可解释性'的多重含义,解析其术语演变与学术分野。
Mechanistic?
- 梳理'机制可解释性'在技术、文化等四类用法中的差异。
- 揭示该术语从狭义因果推断到广义解释范式的语义漂移。
- 适合关注模型解释理论演进的研究者与跨领域对话参与者。
随着对神经网络模型(尤其是语言模型)理解需求的增长,'机制可解释性'一词日益流行,但也带来了混淆。本文描述了该术语在可解释性研究中的四种使用方式:最严格的定义要求因果性主张;较宽泛的技术定义允许探索模型内部机制;还有两种文化层面的定义,分别指向特定学术社群及其扩展认知。文章追溯自然语言处理可解释性社区的历史,分析了独立发展的'机制可解释性'社群的形成过程。最后讨论该术语如何被整个可解释性领域接纳,并指出其多义性源于该领域内部的关键分歧。
原文摘要 · Abstract (English)
The rise of the term "mechanistic interpretability" has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has also led to a fair amount of confusion. So, what does it mean to be "mechanistic"? We describe four uses of the term in interpretability research. The most narrow technical definition requires a claim of causality, while a broader technical definition allows for any exploration of a model's internals. However, the term also has a narrow cultural definition describing a cultural movement. To understand this semantic drift, we present a history of the NLP interpretability community and the formation of the separate, parallel "mechanistic" interpretability community. Finally, we discuss the broad cultural definition -- encompassing the entire field of interpretability -- and why the traditional NLP interpretability community has come to embrace it. We argue that the polysemy of "mechanistic" is the product of a critical divide within the interpretability community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。