用因子图生成可理解解释,对抗未知干扰更可靠
Factor Graph-based Interpretable Neural Networks
- 用因子图建模解释与类别间的逻辑规则
- 在推理时识别并修正违反现实逻辑的错误解释
- 无需训练即可应对未知攻击,适合安全敏感场景
可理解的神经网络解释是深入理解决策过程的基础,尤其在输入数据被恶意扰动时更为关键。现有方法通常通过对抗训练缓解扰动影响,但在未知扰动下无法生成可理解解释。为此,我们提出AGAIN——一种基于因子图的可解释神经网络,在未知扰动下仍能生成合理解释。不同于以往需重训练的方法,AGAIN在推理阶段直接整合逻辑规则,通过因子图识别并修正解释中的逻辑错误。我们构建因子图以表达解释与类别间的逻辑关系,并将逻辑规则作为外生知识,使模型能检测违背现实逻辑的不一致解释。进一步提出交互式干预开关策略,基于因子图的逻辑引导修正解释,无需学习扰动模式,突破了对抗训练仅能防御已知扰动的局限性。理论证明显示,解释的可理解性与因子图密切相关。在三个数据集上的大量实验表明,AGAIN性能优于当前最优基线。
原文摘要 · Abstract (English)
Comprehensible neural network explanations are foundations for a better understanding of decisions, especially when the input data are infused with malicious perturbations. Existing solutions generally mitigate the impact of perturbations through adversarial training, yet they fail to generate comprehensible explanations under unknown perturbations. To address this challenge, we propose AGAIN, a fActor GrAph-based Interpretable neural Network, which is capable of generating comprehensible explanations under unknown perturbations. Instead of retraining like previous solutions, the proposed AGAIN directly integrates logical rules by which logical errors in explanations are identified and rectified during inference. Specifically, we construct the factor graph to express logical rules between explanations and categories. By treating logical rules as exogenous knowledge, AGAIN can identify incomprehensible explanations that violate real-world logic. Furthermore, we propose an interactive intervention switch strategy rectifying explanations based on the logical guidance from the factor graph without learning perturbations, which overcomes the inherent limitation of adversarial training-based methods in defending only against known perturbations. Additionally, we theoretically demonstrate the effectiveness of employing factor graph by proving that the comprehensibility of explanations is strongly correlated with factor graph. Extensive experiments are conducted on three datasets and experimental results illustrate the superior performance of AGAIN compared to state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。