提出可做统计推断的特征交互重要性度量方法
iLOCO: Distribution-Free Inference for Feature Interactions
- 基于留一协变量法,量化特征间配对或高阶交互影响
- 无需分布假设,直接给出交互重要性的置信区间
- 结合集成学习提升效率,适合真实数据验证
特征重要性度量在理解模型行为、指导特征选择和提升可解释性方面至关重要。然而,许多机器学习模型涉及复杂特征交互,现有重要性指标难以捕捉配对或高阶效应,而现有交互度量方法普遍存在适用性有限或计算成本过高问题,且缺乏对特征交互进行统计推断的方法。为此,本文首次提出一种模型无关的交互留一协变量法(iLOCO),用于衡量配对特征交互的重要性,并可扩展至高阶交互。进一步,利用最新的LOCO推断进展,构建了无需分布假设、假设极少的iLOCO置信区间。为解决计算挑战,还引入集成学习方法高效计算iLOCO及置信区间,兼具计算与统计效率。在合成数据和真实数据集上的验证表明,该方法优于现有方法,首次实现了特征交互的统计推断能力。
原文摘要 · Abstract (English)
Feature importance measures are widely studied and are essential for understanding model behavior, guiding feature selection, and enhancing interpretability. However, many machine learning fitted models involve complex interactions between features. Existing feature importance metrics fail to capture these pairwise or higher-order effects, while existing interaction metrics often suffer from limited applicability or excessive computation; no methods exist to conduct statistical inference for feature interactions. To bridge this gap, we first propose a new model-agnostic metric, interaction Leave-One-Covariate-Out (iLOCO), for measuring the importance of pairwise feature interactions, with extensions to higher-order interactions. Next, we leverage recent advances in LOCO inference to develop distribution-free and assumption-light confidence intervals for our iLOCO metric. To address computational challenges, we also introduce an ensemble learning method for calculating the iLOCO metric and confidence intervals that we show is both computationally and statistically efficient. We validate our iLOCO metric and our confidence intervals on both synthetic and real data sets, showing that our approach outperforms existing methods and provides the first inferential approach to detecting feature interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。