arXiv:2603.05024cs.AIcs.LG2026-03

提出解释稳定性度量CIES,评估商业模型解释在噪声下的可信度。

Measuring the Fragility of Trust: Devising Credibility Index via Explanation Stability (CIES) for Business Decision Support Systems

  • 用加权距离函数衡量解释在真实业务噪声下的稳定性,突出关键特征变化的影响。
  • 实验显示模型复杂度和数据平衡方式显著影响解释可信度,CIES区分能力更强(p<0.01)。
  • 适合关注AI决策可信度的业务分析师、风控人员及模型部署者使用。

可解释人工智能(XAI)方法(如SHAP、LIME)在高风险商业场景中日益普及,但其解释的可信度及其在现实数据扰动下的稳定性尚未量化。本文提出基于解释稳定性的可信度指数(CIES),一种数学严谨的度量方法,用于评估模型解释在真实业务噪声下的鲁棒性。CIES关注预测背后的理由是否一致,而不仅是预测本身。该度量采用秩加权距离函数,对重要特征的不稳定性给予更大惩罚,体现商业语义——核心决策因素的变化比边缘特征更关键。我们在三个数据集(客户流失、信用风险、员工流失)、四种树基分类模型及两种数据平衡条件(含SMOTE)下评估CIES。结果表明:模型复杂度影响解释可信度,使用SMOTE处理类别不平衡不仅改变预测性能,也影响解释稳定性;与均匀基线度量相比,CIES在全部24种配置下均表现出统计显著更强的区分能力(p < 0.01)。四类噪声水平下的敏感性分析验证了该度量本身的稳健性。研究为业务实践提供了可落地的‘可信度预警系统’。

原文摘要 · Abstract (English)

Explainable Artificial Intelligence (XAI) methods (SHAP, LIME) are increasingly adopted to interpret models in high-stakes businesses. However, the credibility of these explanations, their stability under realistic data perturbations, remains unquantified. This paper introduces the Credibility Index via Explanation Stability (CIES), a mathematically grounded metric that measures how robust a model's explanations are when subject to realistic business noise. CIES captures whether the reasons behind a prediction remain consistent, not just the prediction itself. The metric employs a rank-weighted distance function that penalizes instability in the most important features disproportionately, reflecting business semantics where changes in top decision drivers are more consequential than changes in marginal features. We evaluate CIES across three datasets (customer churn, credit risk, employee attrition), four tree-based classification models and two data balancing conditions. Results demonstrate that model complexity impacts explanation credibility, class imbalance treatment via SMOTE affects not only predictive performance but also explanation stability, and CIES provides statistically superior discriminative power compared to a uniform baseline metric (p < 0.01 in all 24 configurations). A sensitivity analysis across four noise levels confirms the robustness of the metric itself. These findings offer business practitioners a deployable "credibility warning system" for AI-driven decision support.

可解释AI可信度评估业务决策解释稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。