综述病理学多模态基础模型,整合图像与文本、基因等数据提升诊断能力。
Multi-Modal Foundation Models for Computational Pathology: A Survey
- 按视觉-文本、视觉-知识图谱、视觉-基因表达分三类模型
- 梳理32个前沿模型和28个专用多模态数据集
- 适合关注医学AI融合的科研人员和临床开发者
基础模型已成为计算病理学(CPath)的重要范式,可实现组织病理图像的可扩展、通用分析。早期研究聚焦仅基于视觉数据的单模态模型,近年进展表明融合文本报告、结构化知识与分子谱等异构数据的多模态基础模型具有巨大潜力。本文系统综述了当前多模态基础模型在CPath中的发展,重点覆盖基于苏木精-伊红(H&E)染色全切片图像(WSI)和块级表示的模型。我们将32个前沿多模态基础模型归为三大范式:视觉-语言、视觉-知识图谱、视觉-基因表达,并将视觉-语言模型进一步分为非大语言模型(LLM)与基于LLM的方法。此外,我们分析了28个专用于病理学的多模态数据集,按图像-文本对、指令数据集及图像-其他模态对分类。本综述还构建了下游任务分类体系,总结训练与评估策略,并指出关键挑战与未来方向。旨在为病理与AI交叉领域的研究人员提供重要参考。
原文摘要 · Abstract (English)
Foundation models have emerged as a powerful paradigm in computational pathology (CPath), enabling scalable and generalizable analysis of histopathological images. While early developments centered on uni-modal models trained solely on visual data, recent advances have highlighted the promise of multi-modal foundation models that integrate heterogeneous data sources such as textual reports, structured domain knowledge, and molecular profiles. In this survey, we provide a comprehensive and up-to-date review of multi-modal foundation models in CPath, with a particular focus on models built upon hematoxylin and eosin (H&E) stained whole slide images (WSIs) and tile-level representations. We categorize 32 state-of-the-art multi-modal foundation models into three major paradigms: vision-language, vision-knowledge graph, and vision-gene expression. We further divide vision-language models into non-LLM-based and LLM-based approaches. Additionally, we analyze 28 available multi-modal datasets tailored for pathology, grouped into image-text pairs, instruction datasets, and image-other modality pairs. Our survey also presents a taxonomy of downstream tasks, highlights training and evaluation strategies, and identifies key challenges and future directions. We aim for this survey to serve as a valuable resource for researchers and practitioners working at the intersection of pathology and AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。