arXiv:2411.19666eess.IVcs.AI2024-11被引 245

TITAN用33万张病理切片训练,能无标签生成报告并识别罕见癌

Multimodal Whole Slide Foundation Model for Pathology

  • 用33万张切片+42万自动生成报告预训练,实现跨模态对齐
  • 零样本/少样本下准确率超现有模型,罕见癌检索效果显著
  • 适合资源匮乏场景,如稀有病诊断与癌症预后分析

计算病理学因基础模型的发展而革新,这些模型通过自监督学习将组织病理学感兴趣区域(ROIs)编码为通用且可迁移的特征表示。然而,由于特定疾病队列中临床数据有限,尤其是在罕见疾病方面,将这些进展应用于患者级和切片级复杂临床挑战仍受限制。我们提出TITAN,一种多模态全切片基础模型,基于335,645张全切片图像(WSIs)通过视觉自监督学习及与相应病理报告的视觉-语言对齐进行预训练,并利用多模态生成式AI助手生成423,122条合成描述。无需微调或临床标签,TITAN即可提取通用切片表示并生成泛化性强的病理报告,适用于资源有限的临床场景,如罕见疾病检索和癌症预后预测。我们在多种临床任务上评估TITAN,发现其在线性探测、少样本与零样本分类、罕见癌症检索、跨模态检索及病理报告生成等任务中,均优于现有的ROI与切片基础模型。

原文摘要 · Abstract (English)

The field of computational pathology has been transformed with recent advances in foundation models that encode histopathology region-of-interests (ROIs) into versatile and transferable feature representations via self-supervised learning (SSL). However, translating these advancements to address complex clinical challenges at the patient and slide level remains constrained by limited clinical data in disease-specific cohorts, especially for rare clinical conditions. We propose TITAN, a multimodal whole slide foundation model pretrained using 335,645 WSIs via visual self-supervised learning and vision-language alignment with corresponding pathology reports and 423,122 synthetic captions generated from a multimodal generative AI copilot for pathology. Without any finetuning or requiring clinical labels, TITAN can extract general-purpose slide representations and generate pathology reports that generalize to resource-limited clinical scenarios such as rare disease retrieval and cancer prognosis. We evaluate TITAN on diverse clinical tasks and find that TITAN outperforms both ROI and slide foundation models across machine learning settings such as linear probing, few-shot and zero-shot classification, rare cancer retrieval and cross-modal retrieval, and pathology report generation.

病理分析多模态生成模型罕见病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。