arXiv:2608.23817cs.AIcs.CY2026-08中稿 · publication in the…

为可解释AI提供可信度审计框架,评估解释的稳定性与准确性。

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

论文配图:A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification
图 1 · 摘自论文原文
  • 构建基于扰动测试的审计协议,量化解释的鲁棒性与保真度。
  • 实测显示高精度模型仍可能生成无效解释,过拟合导致保真度失效。
  • 适用于医疗、粮食安全等关键领域,确保AI决策可信赖。

SHAP和LIME是当前解释黑箱模型的标准工具,但在输入受微小噪声扰动时,其输出可能显著变化——我们在马达加斯加粮食安全研究中已观察到此问题(Ralinirina et al., 2025)。这种不稳定性引发质疑:这些解释是否可信?我们通过构建审计协议,衡量任意后处理解释器的两个属性:鲁棒性(输入扰动下解释的稳定性)和保真度(被识别为重要特征是否真实驱动预测)。两者合并为单一可信度分数。在马达加斯加多部门数据集(83个特征,253条记录,4类营养不良)上,使用三种分类器和两种解释器及其正则化版本进行测试。结果令人警醒:AUC超过0.99的模型仍可能产生数值退化或完全无信息的解释,且当模型过拟合时,保真度得分失去区分能力。这表明,在敏感领域中,对XAI输出进行审计并非可选项,而是必要环节。

原文摘要 · Abstract (English)

SHAP and LIME are now standard tools for interpreting black-box predictions, yet their outputs can vary substantially when the input is perturbed by small amounts of noise--a problem we observed firsthand in our previous work on food security in Madagascar (Ralinirina et al., 2025). This variability raises the question of whether such explanations can be trusted at all. We address it by constructing an auditing protocol that measures two properties of any post-hoc explainer: robustness (how stable the explanation is under input perturbation) and fidelity (whether the features deemed important actually drive the model's prediction). These two quantities are combined into a single Trust Score. We run the protocol on a multi-sectoral dataset from Madagascar (83 features, 253 records, 4 malnutrition classes) using three classifiers and two explainers, plus their regularized counterparts. The results are sobering: models with AUC above 0.99 can produce numerically degenerate or flatly uninformative explanations, and fidelity scores lose discriminative power when the model is overfitted. These findings suggest that auditing XAI outputs is not optional but necessary, particularly when they inform decisions in sensitive domains.

可解释AI模型审计信任度鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。