arXiv:2508.09639cs.AI2025-08被引 7

为SHAP值提供不确定性分解,提升高风险场景下的解释可靠性。

UbiQTree: Uncertainty Quantification in XAI with Tree Ensembles

  • 用证据理论与狄利克雷过程对树集成模型的SHAP值做不确定性分解。
  • 发现高SHAP值特征未必稳定,其不确定性可随数据质量提升而降低。
  • 适合医疗等高风险领域中需可信解释的模型评估与优化场景。

解释性人工智能(XAI)技术如SHAP已成为解释复杂树集成模型的重要工具,尤其在医疗分析等高风险领域。然而,传统SHAP值通常作为点估计,忽略了预测模型和数据中的固有不确定性,主要来自可变性(aleatoric)和知识不足(epistemic)。本文提出一种将SHAP值不确定性分解为可变性、知识不足及纠缠成分的方法,结合德普斯特-谢弗证据理论与狄利克雷过程的假设采样,应用于树集成模型。通过三个真实世界案例的统计分析,揭示了嵌入在SHAP解释中的知识不足特性。实验表明,高SHAP值特征并不一定具备稳定性,而通过更优质、更具代表性的数据及合理建模策略,可有效降低知识不足不确定性。树模型,尤其是袋装(bagging)方法,有助于高效量化知识不足不确定性,从而增强解释的可靠性和可解释性,支持高风险应用中的稳健决策与模型改进。

原文摘要 · Abstract (English)

Explainable Artificial Intelligence (XAI) techniques, such as SHapley Additive exPlanations (SHAP), have become essential tools for interpreting complex ensemble tree-based models, especially in high-stakes domains such as healthcare analytics. However, SHAP values are usually treated as point estimates, which disregards the inherent and ubiquitous uncertainty in predictive models and data. This uncertainty has two primary sources: aleatoric and epistemic. The aleatoric uncertainty, which reflects the irreducible noise in the data. The epistemic uncertainty, which arises from a lack of data. In this work, we propose an approach for decomposing uncertainty in SHAP values into aleatoric, epistemic, and entanglement components. This approach integrates Dempster-Shafer evidence theory and hypothesis sampling via Dirichlet processes over tree ensembles. We validate the method across three real-world use cases with descriptive statistical analyses that provide insight into the nature of epistemic uncertainty embedded in SHAP explanations. The experimentations enable to provide more comprehensive understanding of the reliability and interpretability of SHAP-based attributions. This understanding can guide the development of robust decision-making processes and the refinement of models in high-stakes applications. Through our experiments with multiple datasets, we concluded that features with the highest SHAP values are not necessarily the most stable. This epistemic uncertainty can be reduced through better, more representative data and following appropriate or case-desired model development techniques. Tree-based models, especially bagging, facilitate the effective quantification of epistemic uncertainty.

XAI不确定性量化SHAP树集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。