arXiv:2506.08505cs.LGcs.AI2025-06ICML被引 14

提出高效生成可证明神经网络解释的新方法

Explaining, Fast and Slow: Abstraction and Refinement of Provable Explanations

  • 用简化网络抽象原模型,加速可证明解释计算
  • 在多个抽象层级上获得精细预测解释
  • 适合需要可信解释的高安全性场景

尽管后验可解释性技术取得显著进展,但多数方法依赖启发式策略,无法提供解释的正式保证。近期研究发现,通过神经网络验证技术识别使预测不变的输入特征子集,可获得具有形式保证的解释。然而,此类解释的计算面临严重可扩展性挑战。本文提出一种新颖的抽象-精化技术,高效计算神经网络预测的可证明充分解释。方法通过构建大幅简化后的网络来抽象原始大模型,其充分解释对原模型也具可证明充分性,从而显著加速验证过程。若简化网络的解释不足,则逐步增大网络规模直至收敛。实验表明,该方法在提升可证明解释效率的同时,还能在不同抽象层级上提供对网络预测的细粒度解释。

原文摘要 · Abstract (English)

Despite significant advancements in post-hoc explainability techniques for neural networks, many current methods rely on heuristics and do not provide formally provable guarantees over the explanations provided. Recent work has shown that it is possible to obtain explanations with formal guarantees by identifying subsets of input features that are sufficient to determine that predictions remain unchanged using neural network verification techniques. Despite the appeal of these explanations, their computation faces significant scalability challenges. In this work, we address this gap by proposing a novel abstraction-refinement technique for efficiently computing provably sufficient explanations of neural network predictions. Our method abstracts the original large neural network by constructing a substantially reduced network, where a sufficient explanation of the reduced network is also provably sufficient for the original network, hence significantly speeding up the verification process. If the explanation is in sufficient on the reduced network, we iteratively refine the network size by gradually increasing it until convergence. Our experiments demonstrate that our approach enhances the efficiency of obtaining provably sufficient explanations for neural network predictions while additionally providing a fine-grained interpretation of the network's predictions across different abstraction levels.

可解释性神经网络形式验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。