arXiv:2602.11760stat.MLcs.LG2026-02被引 1

用集成模型提升特征重要性估计的稳定性,尤其适合复杂生物医学数据。

Aggregate Models, Not Explanations: Improving Feature Importance Estimation

  • 在模型层面集成,而非聚合解释结果,能更准确估计特征重要性。
  • 针对高表达力模型,该方法显著降低变量重要性估计误差。
  • 适用于需要可靠特征分析的生物医学研究,如蛋白质组学分析。

特征重要性方法有望将机器学习模型从预测工具转变为科学发现工具。然而,由于数据采样和算法随机性,高表达力模型可能不稳定,导致特征重要性估计不准确,影响其在关键生物医学应用中的实用性。尽管集成学习可缓解此问题,但选择对单个集成模型进行解释还是聚合各独立模型的解释,因重要性度量的非线性而难以决策,且研究不足。我们的理论分析在支持复杂先进机器学习模型的假设下表明,这一选择主要由模型的超额风险决定。与以往研究不同,我们证明在模型层面集成能更准确估计变量重要性,尤其对高表达力模型,通过减少主导误差项实现。我们在经典基准和来自英国生物样本库的大规模蛋白质组学研究中验证了这些发现。

原文摘要 · Abstract (English)

Feature-importance methods show promise in transforming machine learning models from predictive engines into tools for scientific discovery. However, due to data sampling and algorithmic stochasticity, expressive models can be unstable, leading to inaccurate variable importance estimates and undermining their utility in critical biomedical applications. Although ensembling offers a solution, deciding whether to explain a single ensemble model or aggregate individual model explanations is difficult due to the nonlinearity of importance measures and remains largely understudied. Our theoretical analysis, developed under assumptions accommodating complex state-of-the-art ML models, reveals that this choice is primarily driven by the model's excess risk. In contrast to prior literature, we show that ensembling at the model level provides more accurate variable-importance estimates, particularly for expressive models, by reducing this leading error term. We validate these findings on classical benchmarks and a large-scale proteomic study from the UK Biobank.

特征重要性集成学习生物医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。