arXiv:2605.15154stat.MLcs.LG2026-05

提出鲁棒特征归因方法RoSHAP,让模型解释更稳定可靠。

RoSHAP: A Distributional Framework and Robust Metric for Stable Feature Attribution

论文配图:RoSHAP: A Distributional Framework and Robust Metric for Stable Feature Attribution
图 1 · 摘自论文原文
  • 基于分布建模和自助采样,量化特征归因的随机性。
  • 在模拟与真实数据中,比传统方法更准识别关键特征。
  • 适合需要稳定解释的高可靠性场景,如医疗、金融决策。

特征归因分析对理解机器学习模型和支撑可靠数据决策至关重要。然而,特征归因值常因训练-测试划分、随机种子或模型拟合过程不同而产生显著波动,导致归因结果不稳定。本文提出一个融合归因随机性的框架,并引入鲁棒归因度量RoSHAP,基于SHAP实现稳定特征排序。该框架通过自助采样与核密度估计建模归因分数分布;在温和正则条件下,聚合归因分数渐近服从高斯分布,大幅降低分布估计计算成本。RoSHAP将SHAP分布归纳为综合反映特征活跃性、强度与稳定性的鲁棒排序准则。仿真与真实数据实验表明,该框架及RoSHAP在识别信号特征方面优于标准单次运行归因方法。使用RoSHAP筛选特征构建的模型,在仅用少量预测变量的情况下,仍能达到全特征模型的预测性能。该方法提升了模型解释的稳定性与可读性,支持分析中获得一致可靠的洞察。

原文摘要 · Abstract (English)

Feature attribution analysis is critical for interpreting machine learning models and supporting reliable data-driven decisions. However, feature attribution measures often exhibit stochastic variation: different train--test splits, random seeds, or model-fitting procedures can produce substantially different attribution values and feature rankings. This paper proposes a framework for incorporating stochastic nature of feature attribution and a robust attribution metric, RoSHAP, for stable feature ranking based on the SHAP metric. The proposed framework models the distribution of feature attribution scores and estimates it through bootstrap resampling and kernel density estimation. We show that, under mild regularity conditions, the aggregated feature attribution score is asymptotically Gaussian, which greatly reduces the computational cost of distribution estimation. The RoSHAP summarizes the distribution of SHAP into a robust feature-ranking criterion that simultaneously rewards features that are active, strong, and stable. Through simulations and real-data experiments, the proposed framework and RoSHAP outperform standard single-run attribution measures in identifying signal features. In addition, models built using RoSHAP-selected features achieve predictive performance comparable to full-feature models while using substantially fewer predictors. The proposed RoSHAP approach improves the stability and interpretability of machine learning models, enabling reliable and consistent insights for analysis.

特征归因模型可解释性稳定性优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。