提出MUCE方法,用图形和指标揭示模型预测的局部变化与可信度。
Explaining Machine Learning Predictive Models through Conditional Expectation Methods
- 基于多变量条件期望,通过邻近采样捕捉特征交互对预测的影响。
- 在合成与真实数据上验证,能有效识别决策边界附近的复杂行为。
- 新增稳定性和不确定性指标,支持可视化与量化双重解释。
复杂机器学习模型因难以解释而被称为黑箱,影响用户理解与信任,尤其在高风险场景中。本文提出一种无需依赖模型结构的局部可解释性方法——多变量条件期望(MUCE),扩展了个体条件期望(ICE)方法,在推理时对给定样本邻域进行多维网格采样,生成展示预测变化趋势的图像。同时引入稳定性与不确定性两个量化指标,进一步分解为不确定性+与不确定性-,以捕捉全局方法易忽略的非对称效应。在三个数据集上评估:两个二维与三维合成数据用于测试决策边界附近的行为,一个经过变换的真实世界数据集用于检验对异质特征类型的适应性。结果表明,MUCE能有效揭示复杂局部行为,指标也提供了预测置信度的有意义洞察。该方法结合改进的ICE与新指标,为局部可解释性提供了实用工具,兼具图形化与定量分析能力,提升模型透明度与可信度。
原文摘要 · Abstract (English)
The rapid adoption of complex Artificial Intelligence (AI) and Machine Learning (ML) models has led to their characterization as black boxes due to the difficulty of explaining their internal decision-making processes. This lack of transparency hinders users' ability to understand, validate and trust model behavior, particularly in high-risk applications. Although explainable AI (XAI) has made significant progress, there remains a need for versatile and effective techniques to address increasingly complex models. This work introduces Multivariate Conditional Expectation (MUCE), a model-agnostic method for local explainability designed to capture prediction changes from feature interactions. MUCE extends Individual Conditional Expectation (ICE) by exploring a multivariate grid of values in the neighborhood of a given observation at inference time, providing graphical explanations that illustrate the local evolution of model predictions. In addition, two quantitative indices, stability and uncertainty, summarize local behavior and assess model reliability. Uncertainty is further decomposed into uncertainty+ and uncertainty- to capture asymmetric effects that global measures may overlook. The proposed method is validated using XGBoost models trained on three datasets: two synthetic (2D and 3D) to evaluate behavior near decision boundaries, and one transformed real-world dataset to test adaptability to heterogeneous feature types. Results show that MUCE effectively captures complex local model behavior, while the stability and uncertainty indices provide meaningful insight into prediction confidence. MUCE, together with the ICE modification and the proposed indices, offers a practical contribution to local explainability, enabling both graphical and quantitative insights that enhance the interpretability of predictive models and support more trustworthy and transparent decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。