arXiv:2505.08198stat.MLcs.LG2025-05被引 1

提出高效稳定的谢帕利值近似方法,大幅降低计算成本。

SIM-Shapley: A Stable and Computationally Efficient Approach to Shapley Value Approximation

  • 基于随机优化思想设计迭代算法,提升稳定性与效率。
  • 实测计算时间减少最高达85%,特征归因质量接近顶尖方法。
  • 适用于高维场景,适合需要可信解释的医疗金融领域。

可解释人工智能(XAI)对可信机器学习至关重要,尤其在医疗、金融等高风险领域。谢帕利值(SV)方法为复杂模型提供严谨的特征归因框架,但计算成本高昂,限制其在高维场景的扩展性。本文提出一种受随机优化启发的稳定高效近似方法——随机迭代动量谢帕利值(SIM-Shapley)。理论上分析方差,证明线性 $Q$-收敛性,并在真实数据集上验证其出色的实践稳定性和低偏差。数值实验表明,相比现有最优基线,SIM-Shapley 最多可减少 85% 的计算时间,同时保持相近的特征归因质量。该方法的随机小批量迭代框架还可推广至更广泛的样本平均近似问题,为兼具计算效率与稳定性的优化提供新路径。代码已公开于 https://github.com/nliulab/SIM-Shapley。

原文摘要 · Abstract (English)

Explainable artificial intelligence (XAI) is essential for trustworthy machine learning (ML), particularly in high-stakes domains such as healthcare and finance. Shapley value (SV) methods provide a principled framework for feature attribution in complex models but incur high computational costs, limiting their scalability in high-dimensional settings. We propose Stochastic Iterative Momentum for Shapley Value Approximation (SIM-Shapley), a stable and efficient SV approximation method inspired by stochastic optimization. We analyze variance theoretically, prove linear $Q$-convergence, and demonstrate improved empirical stability and low bias in practice on real-world datasets. In our numerical experiments, SIM-Shapley reduces computation time by up to 85% relative to state-of-the-art baselines while maintaining comparable feature attribution quality. Beyond feature attribution, our stochastic mini-batch iterative framework extends naturally to a broader class of sample average approximation problems, offering a new avenue for improving computational efficiency with stability guarantees. Code is publicly available at https://github.com/nliulab/SIM-Shapley.

可解释AI特征归因高效算法随机优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。