arXiv:2604.15107stat.MLcs.LG2026-04

提出新方法MinShap,精准识别冗余特征,提升模型解释性。

MinShap: A Shapley-Based Framework for Feature Redundancy

  • 用最小值聚合替代平均,检测特征在所有条件下是否仍重要
  • 在假设下精确识别冗余特征,理论证明优于现有方法
  • 适用于需要稳定特征选择的机器学习场景

Shapley值虽能灵活分配特征贡献,但不天然适合特征选择:即使某特征在其余变量条件下冗余,仍可能获得正贡献。本文提出MinShap框架,通过条件重要性函数VI_j^S识别重要或非冗余特征。不同于传统平均聚合,MinShap采用最小值聚合,检验特征在每种条件下的必要性。在简单零单调性假设下,最小值聚合可精确刻画特征冗余,提供严谨的特征选择准则。该方法统一了统计特征选择与基于表示的可解释性,同时保留了Shapley聚合的稳定性。我们开发了具备统计保证的可扩展算法,建立与多重检验程序的联系,并通过理论和实验表明,MinShap在准确性和稳定性上优于现有无模型方法。

原文摘要 · Abstract (English)

Shapley values provide a flexible framework for attributing feature contributions to model predictions, but they are not naturally suited for feature selection: a feature may receive a positive attribution even when it is redundant given the remaining variables. In this paper, we introduce \textbf{MinShap}, a general framework for identifying \emph{important} or \emph{non-redundant} features through conditional importance functionals $VI_j^S$. Rather than averaging feature contributions across conditioning sets, MinShap aggregates them using the \emph{minimum}, thereby testing whether a feature remains relevant under every conditioning context. We show that, under a simple \emph{null monotonicity} condition, the minimum aggregation exactly characterizes feature redundancy and yields a principled feature selection criterion. This perspective provides a unified framework for statistical feature selection and representation-based interpretability while retaining the stability advantages of Shapley-style aggregation. We develop scalable algorithms with statistical guarantees, establish connections to multiple-testing procedures, and demonstrate through theory and experiments that MinShap produces more accurate and stable feature selection than existing model-agnostic approaches.

特征选择可解释性Shapley值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。