提出临床表示学习新方法,提升模型在不同医院的泛化能力
Learning Clinical Representations Under Systematic Distribution Shift
- 分离生理信号与机构特异性记录习惯,学习不变表示
- 跨医院测试中,AUC最高提升2~3点,保持原性能
- 适合医疗数据分布差异大的真实场景应用
临床机器学习模型常基于大规模多模态基础框架训练,但部署环境与训练数据生成条件存在系统性差异,源于测量策略、文档习惯和机构流程的异质性,导致生理信号与机构特异性特征纠缠。本文提出一种面向多模态临床预测的实践无关表示学习框架,将临床观测建模为潜在生理因素与环境依赖过程的产物,引入联合优化目标,在提升预测性能的同时抑制嵌入中环境可预测信息。具体通过监督风险最小化、对抗式环境正则化及跨医院不变风险惩罚实现。在多个纵向电子病历预测任务与跨机构评估中,该方法相比掩码预训练与标准监督基线,分布外AUROC提升2至3点,同时保持分布内性能并改善校准性。结果表明,在表示学习阶段显式处理系统性分布偏移,能获得更鲁棒、可迁移的临床模型,凸显结构不变性对医疗AI的重要性。
原文摘要 · Abstract (English)
Clinical machine learning models are increasingly trained using large scale, multimodal foundation paradigms, yet deployment environments often differ systematically from the data generating settings used during training. Such shifts arise from heterogeneous measurement policies, documentation practices, and institutional workflows, leading to representation entanglement between physiologic signal and practice specific artifacts. In this work, we propose a practice invariant representation learning framework for multimodal clinical prediction. We model clinical observations as arising from latent physiologic factors and environment dependent processes, and introduce an objective that jointly optimizes predictive performance while suppressing environment predictive information in the learned embedding. Concretely, we combine supervised risk minimization with adversarial environment regularization and invariant risk penalties across hospitals. Across multiple longitudinal EHR prediction tasks and cross institution evaluations, our method improves out of distribution AUROC by up to 2 to 3 points relative to masked pretraining and standard supervised baselines, while maintaining in distribution performance and improving calibration. These results demonstrate that explicitly accounting for systematic distribution shift during representation learning yields more robust and transferable clinical models, highlighting the importance of structural invariance alongside architectural scale in healthcare AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。