arXiv:2608.20825cs.AI2026-08

仅靠预测可信度无法确保AI可靠,需结合决策机制验证。

Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress

  • 提出'能力包络'框架,融合预测与解释认证
  • 实证发现预测指标相同的模型解释能力差异巨大
  • 适合关注AI可信性、安全部署的研究者和工程师

人工智能系统在医疗、安防等关键场景中作出重要判断,其可信度常依赖于预测性能指标:准确率、校准度和保真覆盖率。然而这些指标是否足以保障模型可信仍不明确。本文证明:它们不足以建立可信。通过分离定理表明,一个可靠模型与一个被破坏的模型可在所有预测类认证(准确率、校准度、覆盖度)下完全一致,但在解释保真度和实际部署行为上却可任意不同。检测此类失效必须获取模型决策机制信息。本文引入‘能力包络’作为可部署的统一认证框架,结合预测与解释验证。在多种数据集和模型类型中,该框架揭示了仅靠预测认证无法捕捉的失效模式。因此,对预测行为不可见的失效,必须同时考察模型的输出与决策机制。

原文摘要 · Abstract (English)

Artificial intelligence systems increasingly make consequential judgments - which patient is deteriorating, which building is safe to enter, whether an image is authentic and are trusted on the strength of how accurately and confidently they predict. The safeguards that certify them are correspondingly prediction-based: accuracy, calibration and conformal coverage all measure how well a model performs. Whether such checks are sufficient to establish model trustworthiness has remained unclear. Here we prove that they cannot. We establish a separation theorem showing that a reliable model and a compromised one can be identical under every prediction-side certificate, including accuracy, calibration and coverage, yet differ arbitrarily in explanation fidelity and deployment behaviour. Detecting this failure requires access to the model's decision mechanism in addition to its predictions. We introduce the competence envelope as an operational framework that combines prediction and explanation certification into a single deployable criterion. Across diverse datasets and model classes, the proposed framework reveals failure modes that prediction-side certification alone does not capture. Certification against failures that are invisible in prediction behaviour therefore requires evidence about the model's decision mechanism as well as its outputs.

可信AI模型解释认证框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。