无需标注数据,用多模型特征融合生成更通用的病理切片表示
Unsupervised Foundation Model-Agnostic Slide-Level Representation Learning
- 融合多个基础模型的切片特征,通过对比学习构建新表示
- 在4个肿瘤队列上平均提升4.4% AUC,仅用3048张训练切片
- 可兼容未见过的特征提取器,适合跨模型应用
病理全切片图像(WSI)表征学习主要依赖弱监督的多实例学习(MIL),导致表征高度适配特定临床任务。自监督学习(SSL)已成功用于训练组织病理学基础模型(FMs)以生成切片块嵌入,但生成患者或切片级嵌入仍具挑战。现有方法通过对比切片不同增强版本或利用多模态数据将SSL从块级推广至整张切片。本文提出一种新的单模态自监督方法,通过整合多个基础模型的切片块嵌入,在特征空间中生成有效切片表示。所提出的对比预训练策略COBRA结合多个基础模型与基于Mamba-2的架构,在四个公共CPTAC队列上平均性能优于现有最佳切片编码器至少+4.4% AUC,且仅在TCGA的3048张全切片图像上预训练。此外,COBRA在推理时可轻松兼容此前未见的特征提取器。代码开源:https://github.com/KatherLab/COBRA。
原文摘要 · Abstract (English)
Representation learning of pathology whole-slide images (WSIs) has primarily relied on weak supervision with Multiple Instance Learning (MIL). This approach leads to slide representations highly tailored to a specific clinical task. Self-supervised learning (SSL) has been successfully applied to train histopathology foundation models (FMs) for patch embedding generation. However, generating patient or slide level embeddings remains challenging. Existing approaches for slide representation learning extend the principles of SSL from patch level learning to entire slides by aligning different augmentations of the slide or by utilizing multimodal data. By integrating tile embeddings from multiple FMs, we propose a new single modality SSL method in feature space that generates useful slide representations. Our contrastive pretraining strategy, called COBRA, employs multiple FMs and an architecture based on Mamba-2. COBRA exceeds performance of state-of-the-art slide encoders on four different public Clinical Protemic Tumor Analysis Consortium (CPTAC) cohorts on average by at least +4.4% AUC, despite only being pretrained on 3048 WSIs from The Cancer Genome Atlas (TCGA). Additionally, COBRA is readily compatible at inference time with previously unseen feature extractors. Code available at https://github.com/KatherLab/COBRA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。