arXiv:2503.10547cs.CV2025-03

让视觉模型的解释有因果依据,从神经元激活生成可信逻辑规则。

VISIONLOGIC: From Neuron Activations to Causally Grounded Concept Rules for Vision Models

  • 用激活阈值将神经元输出转为逻辑谓词,构建分层解释框架。
  • 通过消融测试验证概念因果性,确保解释与真实决策相关。
  • 在多种模型上生成简洁规则,人类评估显示理解度显著提升。

尽管基于概念的解释比局部归因更具可解释性,但通常依赖相关信号且缺乏因果验证。我们提出VisionLogic,一种新颖的神经符号框架,能生成忠实、分层的全局逻辑规则,其概念经因果验证。该方法首先学习激活阈值,将神经元激活抽象为谓词,再从中推导出类别级逻辑规则;随后通过基于消融的因果测试和迭代区域精炼,将谓词与视觉概念对齐,确保所发现的概念是导致谓词激活的因果因素。在CNN与ViT等多种视觉架构上,该方法生成了可解释的概念与紧凑规则,基本保持原始模型预测性能。大规模人工评估显示,相比以往方法,VisionLogic显著提升了用户对模型行为的理解能力。

原文摘要 · Abstract (English)

While concept-based explanations improve interpretability over local attributions, they often rely on correlational signals and lack causal validation. We introduce VisionLogic, a novel neural-symbolic framework that produces faithful, hierarchical explanations as global logical rules over causally validated concepts. VisionLogic first learns activation thresholds that abstract neuron activations into predicates, then induces class-level logical rules from these predicates. It then grounds predicates to visual concepts via ablation-based causal tests with iterative region refinement, ensuring that discovered concepts correspond to features that are causal for predicate activation. Across different vision architectures such as CNNs and ViTs, it produces interpretable concepts and compact rules that largely preserve the original model's predictive performance. In our large-scale human evaluations, VisionLogic's concept explanations significantly improve participants' understanding of model behavior over prior concept-based methods.

可解释AI因果推理神经符号视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。