arXiv:2608.00989stat.MLcs.LG2026-08

提出可控制组特征错误发现率的新方法,适用于序列与分组模型。

Model-Agnostic FDR Control via Group Gaussian Mirror and Permutation SHAP

  • 用矩阵扰动构造块级镜像统计量,适配分组特征场景。
  • 在相关信号下保持可靠错误发现率控制,提升检测效能。
  • 不依赖模型结构与分布假设,适合各类神经网络和线性模型。

现有错误发现率(FDR)控制的特征选择方法多针对单个特征的坐标假设,无法处理序列或分组模型中一个原始特征由多个子特征构成的情况(如时间滞后、递归状态或注意力交互)。本文提出一种适用于此类场景的分组特征FDR控制框架。对于分组线性模型,采用矩阵值扰动构建零对称的块级镜像统计量;对于神经序列模型,结合置换SHAP导数作为模型无关的块级重要性评分,并辅以核基依赖度量。该框架不依赖特定网络结构,无需指定协变量分布,当块大小为1时退化为高斯镜像或神经高斯镜像。理论证明了低维与高维分组线性模型下的FDR控制,以及固定非线性模型下平滑置换SHAP导数的渐近对称性。模拟与真实数据实验表明,在相关分组特征信号下,该方法能实现可靠的FDR控制并提升检验功效。

原文摘要 · Abstract (English)

Most FDR-controlled feature selection methods are designed for coordinate-wise hypotheses, where each feature has a single weight or importance score. This abstraction fails in sequential and grouped models, where one original feature is represented by a block of sub-features, such as lags, recurrent states, or attention-based interactions. We propose a grouped-feature FDR control framework for such settings. For grouped linear models, we construct null-symmetric block-level mirror statistics with matrix-valued perturbations. For neural sequential models, we combine Permutation SHAP derivatives as model-agnostic block-level importance scores with kernel-based dependence measure. The framework is model-agnostic across network architectures, does not require specifying the covariate distribution, and reduces to Gaussian Mirror or Neural Gaussian Mirror when the block size is one. We prove FDR control for low- and high-dimensional grouped linear models and asymptotic symmetry of smoothed Permutation SHAP derivatives under fixed fitted nonlinear models. Experiments on simulated and real-world datasets show reliable FDR control and improved power under correlated grouped-feature signals.

特征选择FDR控制分组特征SHAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。