预训练的多重实例学习模型能跨器官迁移,提升病理图像分析性能。
Do Multiple Instance Learning Models Transfer?
- 用11种MIL模型在21个预训练任务上评估迁移能力
- 跨器官预训练比从零开始训练效果更好,且在小数据下表现更优
- 适合需要少量标注数据的临床病理研究者使用
多重实例学习(MIL)是计算病理学中从千兆像素组织图像生成临床有意义的切片级嵌入的核心方法。然而,MIL在小规模、弱监督的临床数据集上常表现不佳。与自然语言处理和传统计算机视觉领域广泛使用迁移学习不同,MIL模型的可迁移性仍不清楚。本研究系统评估了预训练MIL模型的迁移学习能力,涵盖11种模型在21个预训练任务上的表现,用于形态学和分子亚型预测。结果表明,即使在不同器官上预训练的MIL模型,也持续优于从零开始训练的模型。此外,在泛癌数据集上预训练可实现跨器官和任务的强泛化能力,优于切片基础模型,且所需预训练数据少得多。这些发现凸显了MIL模型的鲁棒适应性,并证明利用迁移学习可显著提升计算病理学性能。最后,我们提供了一个标准化MIL模型实现和预训练权重收集的资源,可在https://github.com/mahmoodlab/MIL-Lab 获取。
原文摘要 · Abstract (English)
Multiple Instance Learning (MIL) is a cornerstone approach in computational pathology (CPath) for generating clinically meaningful slide-level embeddings from gigapixel tissue images. However, MIL often struggles with small, weakly supervised clinical datasets. In contrast to fields such as NLP and conventional computer vision, where transfer learning is widely used to address data scarcity, the transferability of MIL models remains poorly understood. In this study, we systematically evaluate the transfer learning capabilities of pretrained MIL models by assessing 11 models across 21 pretraining tasks for morphological and molecular subtype prediction. Our results show that pretrained MIL models, even when trained on different organs than the target task, consistently outperform models trained from scratch. Moreover, pretraining on pancancer datasets enables strong generalization across organs and tasks, outperforming slide foundation models while using substantially less pretraining data. These findings highlight the robust adaptability of MIL models and demonstrate the benefits of leveraging transfer learning to boost performance in CPath. Lastly, we provide a resource which standardizes the implementation of MIL models and collection of pretrained model weights on popular CPath tasks, available at https://github.com/mahmoodlab/MIL-Lab
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。