提出可任意组合模态的病理多模态预训练框架,提升诊断准确性
Any-to-Any Learning in Computational Pathology via Triplet Multimodal Pretraining
- 设计灵活的三模态预训练架构,支持任意子集模态输入
- 在生存预测、癌亚型分类等任务上表现优于或接近顶尖模型
- 适合需要处理多源病理数据的研究者和临床辅助系统开发者
计算病理学与人工智能的进步显著提升了对吉比特级全切片图像及基因组等多模态数据的应用。尽管深度学习展现出巨大潜力,仍面临三大挑战:(1) 异构数据融合需复杂策略,简单拼接难以胜任;(2) 常见模态缺失场景要求模型具备强鲁棒性;(3) 下游任务涵盖单模态至多模态,需统一模型支持。为此,我们提出ALTER框架,整合全切片图像(WSIs)、基因组数据与病理报告,采用“任意模态”设计,支持任意子集预训练,实现超越以切片为中心的跨模态表示学习。我们在生存预测、癌症亚型分类、基因突变预测与报告生成等广泛临床任务中评估ALTER,性能优于或媲美当前最优基线。
原文摘要 · Abstract (English)
Recent advances in computational pathology and artificial intelligence have significantly enhanced the utilization of gigapixel whole-slide images and and additional modalities (e.g., genomics) for pathological diagnosis. Although deep learning has demonstrated strong potential in pathology, several key challenges persist: (1) fusing heterogeneous data types requires sophisticated strategies beyond simple concatenation due to high computational costs; (2) common scenarios of missing modalities necessitate flexible strategies that allow the model to learn robustly in the absence of certain modalities; (3) the downstream tasks in CPath are diverse, ranging from unimodal to multimodal, cnecessitating a unified model capable of handling all modalities. To address these challenges, we propose ALTER, an any-to-any tri-modal pretraining framework that integrates WSIs, genomics, and pathology reports. The term "any" emphasizes ALTER's modality-adaptive design, enabling flexible pretraining with any subset of modalities, and its capacity to learn robust, cross-modal representations beyond WSI-centric approaches. We evaluate ALTER across extensive clinical tasks including survival prediction, cancer subtyping, gene mutation prediction, and report generation, achieving superior or comparable performance to state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。