arXiv:2410.09087cs.AIcs.CL2024-10中稿 · EMNLP被引 57

厘清'机制可解释性'的多重含义,解析其术语演变与学术分野。

Mechanistic?

  • 梳理'机制可解释性'在技术、文化等四类用法中的差异。
  • 揭示该术语从狭义因果推断到广义解释范式的语义漂移。
  • 适合关注模型解释理论演进的研究者与跨领域对话参与者。

随着对神经网络模型(尤其是语言模型)理解需求的增长,'机制可解释性'一词日益流行,但也带来了混淆。本文描述了该术语在可解释性研究中的四种使用方式:最严格的定义要求因果性主张;较宽泛的技术定义允许探索模型内部机制;还有两种文化层面的定义,分别指向特定学术社群及其扩展认知。文章追溯自然语言处理可解释性社区的历史,分析了独立发展的'机制可解释性'社群的形成过程。最后讨论该术语如何被整个可解释性领域接纳,并指出其多义性源于该领域内部的关键分歧。

原文摘要 · Abstract (English)

The rise of the term "mechanistic interpretability" has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has also led to a fair amount of confusion. So, what does it mean to be "mechanistic"? We describe four uses of the term in interpretability research. The most narrow technical definition requires a claim of causality, while a broader technical definition allows for any exploration of a model's internals. However, the term also has a narrow cultural definition describing a cultural movement. To understand this semantic drift, we present a history of the NLP interpretability community and the formation of the separate, parallel "mechanistic" interpretability community. Finally, we discuss the broad cultural definition -- encompassing the entire field of interpretability -- and why the traditional NLP interpretability community has come to embrace it. We argue that the polysemy of "mechanistic" is the product of a critical divide within the interpretability community.

可解释性术语辨析语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。