arXiv:2507.21197cs.LGstat.ML2025-07

用可解释性指导医疗模型分群优化,提升预测效果与公平性

An MLI-Guided Framework for Subgroup-Aware Modeling in Electronic Health Records (AdaptHetero)

  • 结合SHAP与聚类,将解释结果转化为分群建模策略
  • 在3个大型EHR数据集上实现最高174.39%的性能提升
  • 适合关注临床模型公平性与可操作性的研究者

机器学习解释(MLI)通常用于增强临床医生信任和提取电子健康记录(EHR)中的洞见,而非指导针对特定亚群体的可操作建模策略。为弥合这一差距,我们提出AdaptHetero——一种新型的基于MLI的框架,将可解释性洞察转化为跨亚群体建模训练与评估的行动指引。在三个大规模EHR数据集(GOSSIS-1-eICU、WiDS、MIMIC-IV)上评估,AdaptHetero持续揭示了在预测重症监护室死亡率、院内死亡率及隐匿性低氧血症时模型行为的异质性。通过整合基于SHAP的解释与无监督聚类,该框架识别出具有临床意义的亚群体特征,在多个亚群中提升预测性能(最高提升174.39%),同时主动预警其他群体的潜在风险。结果表明,该框架在实现更稳健、公平和情境感知的临床部署方面具有巨大潜力。

原文摘要 · Abstract (English)

Machine learning interpretation (MLI) has primarily been leveraged to foster clinician trust and extract insights from electronic health records (EHRs), rather than to guide subgroup-specific, operationalizable modeling strategies. To bridge this gap, we propose AdaptHetero, a novel MLI-driven framework that transforms interpretability insights into actionable guidance for tailoring model training and evaluation across subpopulations. Evaluated on three large-scale EHR datasets -- GOSSIS-1-eICU, WiDS, and MIMIC-IV -- AdaptHetero consistently uncovers heterogeneous model behaviors in predicting ICU mortality, in-hospital death, and hidden hypoxemia. Integrating SHAP-based interpretation with unsupervised clustering, AdaptHetero identifies clinically meaningful, subgroup-specific characteristics, improving predictive performance across many subpopulations (with gains up to 174.39 percent) while proactively flagging potential risks in others. These results highlight the framework's promise for more robust, equitable, and context-aware clinical deployment.

医疗AI可解释性分群建模EHR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。