arXiv:2503.07667cs.LGcs.AI2025-03ICML被引 20

构建大规模多模态临床数据集,推动医疗AI全面评估患者健康。

CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

  • 整合影像、文本、时间序列等多模态临床数据,构建451万患者样本库。
  • 多任务预训练使超声和心电图分析性能提升最高达29%和23%。
  • 适合医疗AI研究者、多模态模型开发者及临床数据平台建设者。

近期临床AI进展在多个领域取得显著突破,但现有基准与模型仍局限于少数模态与任务,限制了大规模多模态方法的发展,难以实现对患者健康状况的全面评估。为此,我们提出临床大规模集成多模态基准CLIMB,统一整合影像、语言、时间序列与图结构等多种临床数据。CLIMB包含451万例患者样本,总数据量达19.01TB,涵盖2D影像、3D视频、时间序列、图结构及多模态数据。通过广泛实证评估,我们发现多任务预训练能显著提升未充分研究领域的性能,超声分析和心电图分析分别较单任务学习提升29%和23%。在CLIMB上预训练也有效增强模型泛化能力,且强单模态编码器性能可经合适融合策略良好迁移至多模态任务。这些发现为新型架构与预训练策略提供基础。代码已开源:https://github.com/DDVD233/climb。

原文摘要 · Abstract (English)

Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modalities and tasks, which hinders the development of large-scale multimodal methods that can make holistic assessments of patient health and well-being. To bridge this gap, we introduce Clinical Large-Scale Integrative Multimodal Benchmark (CLIMB), a comprehensive clinical benchmark unifying diverse clinical data across imaging, language, temporal, and graph modalities. CLIMB comprises 4.51 million patient samples totaling 19.01 terabytes distributed across 2D imaging, 3D video, time series, graphs, and multimodal data. Through extensive empirical evaluation, we demonstrate that multitask pretraining significantly improves performance on understudied domains, achieving up to 29% improvement in ultrasound and 23% in ECG analysis over single-task learning. Pretraining on CLIMB also effectively improves models' generalization capability to new tasks, and strong unimodal encoder performance translates well to multimodal performance when paired with task-appropriate fusion strategies. Our findings provide a foundation for new architecture designs and pretraining strategies to advance clinical AI research. Code is released at https://github.com/DDVD233/climb.

多模态医疗AI数据集预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。