arXiv:2411.09361cs.CVcs.LG2024-11ICLR被引 8

用电子病历时间数据预训练3D医学影像模型,提升疾病风险预测能力。

Time-to-Event Pretraining for 3D Medical Imaging

  • 引入时间到事件预训练,利用纵向电子病历提供长期时间监督信号。
  • 在8个基准任务上平均AUROC提升23.7%,C-index提升29.4%。
  • 适合临床风险预测、医学影像分析与多模态融合研究者。

随着医学基础模型和影像数据的兴起,可扩展的预训练技术为识别未来疾病风险的影像生物标志物提供了新路径。现有自监督方法虽能捕捉器官形态等局部结构特征,但因缺乏时间上下文,难以将像素级标志物与长期健康结果关联。当前方法仅依赖图像和同期文本描述进行监督,无法建模疾病进展。为此,我们提出时间到事件预训练框架,利用大规模配对纵向电子病历(EHR)数据提供时间监督。基于18,945例CT扫描(420万张2D图像)和数千项EHR衍生任务的时间到事件分布,该方法在8个基准任务上平均实现AUROC提升23.7%、Harrell's C-index提升29.4%,且不损害诊断分类性能。本研究为整合纵向EHR与3D影像数据推进临床风险预测奠定基础。

原文摘要 · Abstract (English)

With the rise of medical foundation models and the growing availability of imaging data, scalable pretraining techniques offer a promising way to identify imaging biomarkers predictive of future disease risk. While current self-supervised methods for 3D medical imaging models capture local structural features like organ morphology, they fail to link pixel biomarkers with long-term health outcomes due to a missing context problem. Current approaches lack the temporal context necessary to identify biomarkers correlated with disease progression, as they rely on supervision derived only from images and concurrent text descriptions. To address this, we introduce time-to-event pretraining, a pretraining framework for 3D medical imaging models that leverages large-scale temporal supervision from paired, longitudinal electronic health records (EHRs). Using a dataset of 18,945 CT scans (4.2 million 2D images) and time-to-event distributions across thousands of EHR-derived tasks, our method improves outcome prediction, achieving an average AUROC increase of 23.7% and a 29.4% gain in Harrell's C-index across 8 benchmark tasks. Importantly, these gains are achieved without sacrificing diagnostic classification performance. This study lays the foundation for integrating longitudinal EHR and 3D imaging data to advance clinical risk prediction.

医学影像时间建模预训练风险预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。