arXiv:2506.05184cs.CV2025-06NeurIPS被引 1

用单张显卡让病理模型高效适配临床任务,提升癌症突变预测准确率。

Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis

  • 利用视觉变压器注意力机制融合弱标签切片信息,实现端到端优化。
  • 在膀胱癌和肺腺癌数据上显著优于传统方法,多标签突变分类表现突出。
  • 适合医疗影像研究者在普通设备上部署大型病理模型。

病理基础模型(PFM)已成为分析全切片图像(WSI)的强大工具。然而,由于仅能获得每张切片级别的弱标签,对这些预训练模型进行特定临床任务的适配面临挑战,通常需采用多实例学习(MIL)范式。本文提出一种单显卡任务自适应方法(TAPFM),利用视觉变压器(ViT)注意力进行MIL聚合,并同时优化特征表示与注意力权重。该方法为MIL聚合器与PFM分别维护独立计算图,确保训练过程稳定且与下游任务目标一致。在膀胱癌和肺腺癌的突变预测任务中,跨机构及TCGA队列评估显示,TAPFM持续优于基准方法,其中H-Optimus-0(TAPFM)表现最优。该方法还能有效处理可操作突变的多标签分类。因此,TAPFM使强大预训练模型在标准硬件上实现临床应用成为可能。

原文摘要 · Abstract (English)

Pathology foundation models (PFMs) have emerged as powerful tools for analyzing whole slide images (WSIs). However, adapting these pretrained PFMs for specific clinical tasks presents considerable challenges, primarily due to the availability of only weak (WSI-level) labels for gigapixel images, necessitating multiple instance learning (MIL) paradigm for effective WSI analysis. This paper proposes a novel approach for single-GPU \textbf{T}ask \textbf{A}daptation of \textbf{PFM}s (TAPFM) that uses vision transformer (\vit) attention for MIL aggregation while optimizing both for feature representations and attention weights. The proposed approach maintains separate computational graphs for MIL aggregator and the PFM to create stable training dynamics that align with downstream task objectives during end-to-end adaptation. Evaluated on mutation prediction tasks for bladder cancer and lung adenocarcinoma across institutional and TCGA cohorts, TAPFM consistently outperforms conventional approaches, with H-Optimus-0 (TAPFM) outperforming the benchmarks. TAPFM effectively handles multi-label classification of actionable mutations as well. Thus, TAPFM makes adaptation of powerful pre-trained PFMs practical on standard hardware for various clinical applications.

病理分析视觉变压器多实例学习单卡训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。