arXiv:2411.06367cs.LGcs.AI2024-11被引 1

利用解释不一致性提升模型可靠性,发现并修复潜在问题

BayesNAM: Leveraging Inconsistency for Reliable Explanations

  • 通过贝叶斯神经网络与特征丢弃结合,捕捉解释不一致
  • 实验验证可暴露数据不足或模型结构缺陷
  • 适合关注模型可信度与可解释性改进的研究者

神经加法模型(NAM)是一种新兴的可解释人工智能方法,虽具备高预测性能和直观解释能力,但常出现解释不一致现象。本文指出,这种不一致并非错误,而是多重要特征数据中自然产生的信号。基于简单理论框架,我们证明了不一致性的必然性,并提出贝叶斯神经加法模型(BayesNAM),融合贝叶斯神经网络与特征丢弃技术,理论证明特征丢弃能有效捕获模型不一致性。实验表明,BayesNAM可识别数据不足或模型结构限制等潜在问题,提供更可靠的解释与改进方向。

原文摘要 · Abstract (English)

Neural additive model (NAM) is a recently proposed explainable artificial intelligence (XAI) method that utilizes neural network-based architectures. Given the advantages of neural networks, NAMs provide intuitive explanations for their predictions with high model performance. In this paper, we analyze a critical yet overlooked phenomenon: NAMs often produce inconsistent explanations, even when using the same architecture and dataset. Traditionally, such inconsistencies have been viewed as issues to be resolved. However, we argue instead that these inconsistencies can provide valuable explanations within the given data model. Through a simple theoretical framework, we demonstrate that these inconsistencies are not mere artifacts but emerge naturally in datasets with multiple important features. To effectively leverage this information, we introduce a novel framework, Bayesian Neural Additive Model (BayesNAM), which integrates Bayesian neural networks and feature dropout, with theoretical proof demonstrating that feature dropout effectively captures model inconsistencies. Our experiments demonstrate that BayesNAM effectively reveals potential problems such as insufficient data or structural limitations of the model, providing more reliable explanations and potential remedies.

可解释AI模型可靠性贝叶斯方法解释一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。