arXiv:2508.08724stat.MLcs.LG2025-08

提出一种可控制错误率的分层变量重要性方法,解决医学影像中变量高度相关时的解释难题。

Hierarchical Variable Importance with Statistical Control for Medical Data-Based Prediction

  • 通过分层树结构识别联合预测结果的变量组,提升高相关数据下的解释能力。
  • 在阿尔茨海默病和脑电图数据上成功识别出生物学合理的特征变量。
  • 方法兼具统计严谨性和计算可行性,适合临床研究中的模型可解释性需求。

机器学习在医学影像预测中取得进展,但复杂模型的可解释性仍是瓶颈。现有模型无关的变量重要性方法在处理高度相关的医学数据时效能不足。本文提出分层-条件重要性(Hierarchical-CPI)方法,将变量重要性推断视为发现共同预测结果的变量组问题。通过沿分层树探索子组,保持计算可行性,并实现明确的族系错误率控制。针对高相关性下条件重要性消失的问题,引入基于树的权重分配机制。在两个神经影像数据集上进行基准测试:基于MRI数据分类阿尔茨海默病诊断(ADNI数据集)和分析脑电图数据上的伯杰效应(TDBRAIN数据集),均识别出具有生物学合理性的变量。

原文摘要 · Abstract (English)

Recent advances in machine learning have greatly expanded the repertoire of predictive methods for medical imaging. However, the interpretability of complex models remains a challenge, which limits their utility in medical applications. Recently, model-agnostic methods have been proposed to measure conditional variable importance and accommodate complex non-linear models. However, they often lack power when dealing with highly correlated data, a common problem in medical imaging. We introduce Hierarchical-CPI, a model-agnostic variable importance measure that frames the inference problem as the discovery of groups of variables that are jointly predictive of the outcome. By exploring subgroups along a hierarchical tree, it remains computationally tractable, yet also enjoys explicit family-wise error rate control. Moreover, we address the issue of vanishing conditional importance under high correlation with a tree-based importance allocation mechanism. We benchmarked Hierarchical-CPI against state-of-the-art variable importance methods. Its effectiveness is demonstrated in two neuroimaging datasets: classifying dementia diagnoses from MRI data (ADNI dataset) and analyzing the Berger effect on EEG data (TDBRAIN dataset), identifying biologically plausible variables.

变量重要性医学影像可解释性高相关性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。