arXiv:2608.16773cs.LG2026-08中稿 · KDD

让可解释神经网络支持非欧几里得空间,实现跨架构的严格解释统一。

Beyond $L_2$: Generalizing Abductive Latent Explanations to Diverse Prototype-Based Architectures

  • 将因果潜空间解释框架拓展至球面、高斯等非欧空间
  • 为多种原型模型设计专用距离上界算法,确保解释形式化
  • 首次实现不同可解释模型的严谨可比性,适合关注模型可信性的研究者

原型神经网络因其内置可解释性备受推崇。近期提出的因果潜空间解释(ALE)利用网络内在结构,提供数学保证的预测安全与人类可读解释,依赖于对潜空间距离的紧致边界计算。然而现有ALE方法仅限于欧几里得潜空间,而当前最先进的架构越来越多采用非欧表示(如球面度量、高斯密度、维度投影),导致现有形式化解释方法失效。本文将ALE框架推广至非欧原型架构:针对每种几何变体,系统推导如何映射至已有边界或构建专有边界算法。通过在全训练图像分类器上计算最小子集形式解释验证理论构造。该统一框架首次实现跨架构的可解释性严格比较。

原文摘要 · Abstract (English)

Prototype-based neural networks are hailed as interpretable-by-design architectures. Recently, Abductive Latent Explanations (ALE) were introduced to provide formal, mathematically guaranteed explanations that leverage the intrinsic structure of these networks to ensure both predictive safety and human readability. ALEs rely on computing tight bounds on latent space distances to produce formal explanations. However, existing ALE formulations are rigidly confined to Euclidean latent spaces. This leaves a critical gap: modern state-of-the-art architectures increasingly rely on non-Euclidean representations - such as spherical metrics, Gaussian densities, and dimensional projections - rendering current formal explanation methods incompatible. In this work, we generalize the ALE framework to support non-Euclidean prototype architectures. For each geometric variant, we systematically derive how to either map the architecture to existing bounds or construct novel, architecture-specific bounding algorithms. We validate our theoretical constructions by computing subset-minimal formal explanations on fully trained image classifiers. By unifying these diverse models under a single formal framework, we enable the first rigorous, cross-architecture comparison of their interpretability.

可解释性原型网络非欧空间形式化解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。