arXiv:2512.17316cs.LGcs.AI2025-12

提出可验证的可解释性标准,让模型解释不再靠感觉。

Explanation Beyond Intuition: A Testable Criterion for Inherent Explainability

  • 用图论分解模型结构,生成可验证的局部解释
  • 证明稀疏神经网络可解释而大回归模型不可
  • 为监管提供可落地的合规测试框架

固有可解释性是可解释人工智能(XAI)的金标准,但目前缺乏统一定义和验证方法。现有研究或依赖指标,或诉诸直觉——‘我们一看就懂’。本文提出一个全局适用的固有可解释性判定准则:利用图论表示并分解模型以实现结构化局部解释,并将其重组为全局解释。局部解释以注释形式呈现,构成可验证的假设-证据结构,支持多种解释方法。该准则符合现有对可解释性的直觉认知,解释了为何大型回归模型可能不可解释,而稀疏神经网络却可以。区分了‘可解释’(允许解释)与‘已解释’(已有验证解释)的概念。最后,完整解释了正在新西兰临床使用的冠心病风险预测模型PREDICT,证实其具备固有可解释性。本工作为可解释性研究提供了结构化框架,使监管机构可在合规体系中灵活且严谨地应用此标准。

原文摘要 · Abstract (English)

Inherent explainability is the gold standard in Explainable Artificial Intelligence (XAI). However, there is not a consistent definition or test to demonstrate inherent explainability. Work to date either characterises explainability through metrics, or appeals to intuition - "we know it when we see it". We propose a globally applicable criterion for inherent explainability. The criterion uses graph theory for representing and decomposing models for structure-local explanation, and recomposing them into global explanations. We form the structure-local explanations as annotations, a verifiable hypothesis-evidence structure that allows for a range of explanatory methods to be used. This criterion matches existing intuitions on inherent explainability, and provides justifications why a large regression model may not be explainable but a sparse neural network could be. We differentiate explainable -- a model that allows for explanation -- and \textit{explained} -- one that has a verified explanation. Finally, we provide a full explanation of PREDICT -- a Cox proportional hazards model of cardiovascular disease risk, which is in active clinical use in New Zealand. It follows that PREDICT is inherently explainable. This work provides structure to formalise other work on explainability, and allows regulators a flexible but rigorous test that can be used in compliance frameworks.

可解释性模型验证医疗AI图论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。