arXiv:2505.22199cs.LGcs.AI2025-05ICLR被引 5

用贝叶斯非负决策层提升模型的不确定性估计与可解释性

Enhancing Uncertainty Estimation and Interpretability via Bayesian Non-negative Decision Layer

  • 引入贝叶斯非负决策层,通过稀疏非负潜变量建模复杂依赖关系
  • 在多个数据集上实现更准确的预测与可靠的不确定性估计
  • 适合需要可信推理和可解释性的高风险场景应用

尽管深度神经网络凭借强大的表达能力取得了显著成功,但大多数模型难以满足实际应用中对不确定性估计的需求。同时,深层网络内部特征高度纠缠,导致多种局部解释方法显示多个无关特征共同影响决策,削弱了可解释性。为此,我们提出贝叶斯非负决策层(BNDL),将深度神经网络重新构建成条件贝叶斯非负因子分析模型。通过引入随机潜变量,BNDL 能够建模复杂依赖关系并提供稳健的不确定性估计。潜变量的稀疏性与非负性促使模型学习解耦表示和决策层,从而提升可解释性。我们还提供了理论保证,证明 BNDL 可实现有效的解耦学习。此外,我们设计了一种基于韦布尔变分推断网络的对应变分推断方法,用于近似潜变量后验分布。实验结果表明,得益于更强的解耦能力,BNDL 不仅提升了模型精度,还实现了可靠的不确定性估计和更好的可解释性。

原文摘要 · Abstract (English)

Although deep neural networks have demonstrated significant success due to their powerful expressiveness, most models struggle to meet practical requirements for uncertainty estimation. Concurrently, the entangled nature of deep neural networks leads to a multifaceted problem, where various localized explanation techniques reveal that multiple unrelated features influence the decisions, thereby undermining interpretability. To address these challenges, we develop a Bayesian Non-negative Decision Layer (BNDL), which reformulates deep neural networks as a conditional Bayesian non-negative factor analysis. By leveraging stochastic latent variables, the BNDL can model complex dependencies and provide robust uncertainty estimation. Moreover, the sparsity and non-negativity of the latent variables encourage the model to learn disentangled representations and decision layers, thereby improving interpretability. We also offer theoretical guarantees that BNDL can achieve effective disentangled learning. In addition, we developed a corresponding variational inference method utilizing a Weibull variational inference network to approximate the posterior distribution of the latent variables. Our experimental results demonstrate that with enhanced disentanglement capabilities, BNDL not only improves the model's accuracy but also provides reliable uncertainty estimation and improved interpretability.

不确定性估计可解释性贝叶斯方法解耦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。