arXiv:2602.16823cs.LGcs.LO2026-02被引 8

提出可证明的神经网络电路发现方法,确保结果在连续输入下可靠有效。

Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees

  • 基于神经网络验证技术,自动发现带数学保证的模型内部电路。
  • 在视觉模型上验证,新方法的鲁棒性显著优于传统方法。
  • 适用于需要可信解释的高风险场景,如医疗或自动驾驶。

自动化电路发现是机制可解释性中的核心工具,用于识别神经网络中负责特定行为的内部组件。尽管已有方法取得进展,但通常依赖启发式或近似手段,且无法对连续输入域提供可证明的保证。本文利用神经网络验证的最新进展,提出一套自动化算法,生成具有可证明保证的电路。重点关注三类保证:(1) 输入域鲁棒性,确保电路在连续输入区域内与模型一致;(2) 鲁棒修补,认证在连续修补扰动下的电路对齐性;(3) 最小性,形式化并捕捉多种简洁性概念。我们揭示了这三类保证之间的丰富理论联系,对算法收敛有关键影响。在多个视觉模型上使用最先进的验证器进行实验,结果表明我们的算法生成的电路在鲁棒性方面显著优于标准方法,为可证明的电路发现建立了原则性基础。

原文摘要 · Abstract (English)

*Automated circuit discovery* is a central tool in mechanistic interpretability for identifying the internal components of neural networks responsible for specific behaviors. While prior methods have made significant progress, they typically depend on heuristics or approximations and do not offer provable guarantees over continuous input domains for the resulting circuits. In this work, we leverage recent advances in neural network verification to propose a suite of automated algorithms that yield circuits with *provable guarantees*. We focus on three types of guarantees: (1) *input domain robustness*, ensuring the circuit agrees with the model across a continuous input region; (2) *robust patching*, certifying circuit alignment under continuous patching perturbations; and (3) *minimality*, formalizing and capturing a wide array of various notions of succinctness. Interestingly, we uncover a diverse set of novel theoretical connections among these three families of guarantees, with critical implications for the convergence of our algorithms. Finally, we conduct experiments with state-of-the-art verifiers on various vision models, showing that our algorithms yield circuits with substantially stronger robustness guarantees than standard circuit discovery methods, establishing a principled foundation for provable circuit discovery.

可解释性神经网络验证自动化发现形式化保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。