提出判断模型解释等价性的新方法,让不同模型共享同一算法解释可被验证。
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
- 基于实现等价性定义解释等价,无需显式描述解释内容。
- 开发算法估算解释等价性,适用于Transformer模型。
- 建立解释、电路与表征相似性的理论关联,支持可扩展分析。
机制可解释性(MI)是解释神经网络的新框架。给定任务和模型,MI旨在发现简洁的算法过程——解释——以说明模型在该任务上的决策机制。然而,MI难以规模化与泛化,主要源于两个挑战:缺乏有效的解释定义;且生成解释常为随意过程。本文通过定义并研究解释等价性问题,解决这一难题:在不显式描述解释的情况下,判断两个不同模型是否共享同一解释。核心思想是:若两个解释的所有可能实现均等价,则二者等价。我们提出并形式化该原则,开发估算解释等价性的算法,并在基于Transformer的模型上进行案例研究。为分析该算法,我们引入解释等价性的充要条件,基于模型表示相似性构建。我们的框架同时关联了模型的算法解释、电路结构与表征,为更加严谨的MI评估及自动化、通用化的解释发现方法奠定基础。
原文摘要 · Abstract (English)
Mechanistic interpretability (MI) is an emerging framework for interpreting neural networks. Given a task and model, MI aims to discover a succinct algorithmic process, an interpretation, that explains the model's decision process on that task. However, MI is difficult to scale and generalize. This stems in part from two key challenges: there is no precise notion of a valid interpretation; and, generating interpretations is often an ad hoc process. In this paper, we address these challenges by defining and studying the problem of interpretive equivalence: determining whether two different models share a common interpretation, without requiring an explicit description of what that interpretation is. At the core of our approach, we propose and formalize the principle that two interpretations of a model are equivalent if all of their possible implementations are also equivalent. We develop an algorithm to estimate interpretive equivalence and case study its use on Transformer-based models. To analyze our algorithm, we introduce necessary and sufficient conditions for interpretive equivalence based on models' representation similarity. We provide guarantees that simultaneously relate a model's algorithmic interpretations, circuits, and representations. Our framework lays a foundation for the development of more rigorous evaluation methods of MI and automated, generalizable interpretation discovery methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。