arXiv:2601.04378cs.LGcs.CV2026-01

让神经网络的解释直接生成预测,而非事后编造理由。

Aligned explanations in neural networks

  • 构建逐样本线性模型,使解释与预测同步生成
  • 在图像分类与分割任务中,解释忠实度四项指标全达标
  • 适合需要高可信度解释的医疗、金融等关键领域

随着人工智能在关键决策中日益重要,真正解释神经网络如何做出预测变得至关重要。然而,当前大多数解释方法仅提供事后的合理化说明,无法保证反映模型的真实推理过程。本文提出解释对齐(explanatory alignment)的概念,要求解释直接构建预测,而非事后补足。为在复杂数据域实现此目标,我们提出点式可解释网络(PiNets),一种伪线性架构,能在实例层面形成线性模型。在图像分类与分割任务上的评估表明,PiNets 在四项标准下表现优异:意义性、对齐性、鲁棒性和充分性(MARS)。该工作为融合深度学习的预测能力与线性模型的可解释性提供了原则性基础,推动可信AI与数据驱动科学发现的发展。

原文摘要 · Abstract (English)

As artificial intelligence increasingly drives critical decisions, the ability to genuinely explain how neural networks make predictions is essential for trust. Yet, most current explanation methods offer post-hoc rationalizations rather than guaranteeing a true reflection of the model's reasoning. We introduce the notion of explanatory alignment, a requirement that explanations directly construct predictions rather than rationalize them. To achieve this in complex data domains, we present Pointwise-interpretable Networks (PiNets), a pseudo-linear architecture that forms linear models instance-wise. Evaluated on image classification and segmentation tasks, PiNets demonstrate that their explanations are deeply faithful across four criteria: meaningfulness, alignment, robustness, and sufficiency (MARS). Our contributions pave the way for promising avenues: by reconciling the predictive power of deep learning with the interpretability of linear models, PiNets provide a principled foundation for trustworthy AI and data-driven scientific discovery.

可解释AI神经网络解释对齐线性模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。