arXiv:2507.04362cs.LGphysics.data-an2025-07

用信息论分解特征贡献,精准识别高阶交互作用

Information-theoretic Quantification of High-order Feature Effects in Classification Problems

  • 基于条件互信息与kNN估计,拆解特征的独有、协同与冗余贡献
  • 在真实基因数据中准确捕捉高阶交互,优于传统方法
  • 适合需要理解特征复杂关系的生物医学与模型可解释性研究

理解预测模型中单个特征的贡献仍是可解释机器学习的核心目标。尽管已有多种模型无关的方法估算特征重要性,但往往难以捕捉高阶交互并分离重叠贡献。本文提出一种信息论扩展的高阶特征重要性(Hi-Fi)方法,利用基于k近邻(kNN)的条件互信息(CMI)估计,处理混合离散与连续随机变量。该框架将特征贡献分解为唯一、协同和冗余三部分,提供更丰富的模型无关理解。我们在具有已知高斯结构的合成数据上验证方法,其真值交互模式可通过解析推导获得;进一步在非高斯及来自TCGA-BRCA的真实基因表达数据上测试。结果表明,所提估计器能准确恢复理论预期结果,为基于交互分析的特征选择算法或模型开发提供潜在应用。

原文摘要 · Abstract (English)

Understanding the contribution of individual features in predictive models remains a central goal in interpretable machine learning, and while many model-agnostic methods exist to estimate feature importance, they often fall short in capturing high-order interactions and disentangling overlapping contributions. In this work, we present an information-theoretic extension of the High-order interactions for Feature importance (Hi-Fi) method, leveraging Conditional Mutual Information (CMI) estimated via a k-Nearest Neighbor (kNN) approach working on mixed discrete and continuous random variables. Our framework decomposes feature contributions into unique, synergistic, and redundant components, offering a richer, model-independent understanding of their predictive roles. We validate the method using synthetic datasets with known Gaussian structures, where ground truth interaction patterns are analytically derived, and further test it on non-Gaussian and real-world gene expression data from TCGA-BRCA. Results indicate that the proposed estimator accurately recovers theoretical and expected findings, providing a potential use case for developing feature selection algorithms or model development based on interaction analysis.

特征重要性信息论高阶交互可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。