arXiv:2507.05742eess.IVcs.CV2025-07被引 1

首个整合癌种分型、风险评估与突变预测的病理图像基础模型

Whole Slide Concepts: A Supervised Foundation Model For Pathological Images

  • 基于切片级标签的端到端多任务监督训练,无需大量数据
  • 仅用5%算力即超越自监督模型在七项任务的表现
  • 支持可解释性分析,适合临床辅助诊断与研究者使用

基础模型正在改变计算病理学,为组织病理图像分析提供新方法。然而,传统基础模型通常需要数周时间在大规模数据库上训练,资源消耗巨大。本文提出一种基于全切片图像的监督式端到端多任务学习框架,首次将癌症亚型分类、风险评估和基因突变预测集成于同一模型中。该模型仅需5%的计算资源即可在七个基准任务上超越自监督模型表现。结果表明,监督训练在少量数据下亦可优于自监督,且患者级标签可通过常规临床流程广泛获取,缓解标注难题。此外,模型内置注意力模块提供跨任务可解释性,并可作为未知癌种的肿瘤检测器。为应对闭源数据问题,模型完全在公开数据上训练。代码与模型权重已开源:https://github.com/FraunhoferMEVIS/MedicalMultitaskModeling。

原文摘要 · Abstract (English)

Foundation models (FMs) are transforming computational pathology by offering new ways to analyze histopathology images. However, FMs typically require weeks of training on large databases, making their creation a resource-intensive process. In this paper, we present a training for foundation models from whole slide images using supervised, end-to-end, multitask learning on slide-level labels. Notably, it is the first model to incorporate cancer subtyping, risk estimation, and genetic mutation prediction into one model. The presented model outperforms self-supervised models on seven benchmark tasks while the training only required 5% of the computational resources. The results not only show that supervised training can outperform self-supervision with less data, but also offer a solution to annotation problems, as patient-based labels are widely available through routine clinical processes. Furthermore, an attention module provides a layer of explainability across different tasks and serves as a tumor detector for unseen cancer types. To address the issue of closed-source datasets, the model was fully trained on openly available data. The code and model weights are made available under https://github.com/FraunhoferMEVIS/MedicalMultitaskModeling.

病理图像基础模型多任务学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。