arXiv:2503.17564eess.IVcs.CV2025-03ICCV被引 4

通过多模态适配器实现病理切片模型的统一微调,提升多种癌症预测性能。

ModalTune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-task Learning in Digital Pathology

论文配图:ModalTune: Fine-Tuning Slide-Level Foundation Models with Multi-Modal Information for Multi-task Learning in Digital Pathology
图 1 · 摘自论文原文
  • 引入模态适配器,不修改主模型即可融合新模态数据
  • 在四种癌症上同时提升生存期与亚型预测准确率,达当前最优
  • 首次实现多任务、多模态、泛癌种联合建模,适合临床多病种分析

数字病理中的预测任务因全切片图像(WSIs)规模巨大且训练信号较弱而极具挑战。计算能力提升、数据可用性增加以及自监督学习(SSL)的发展推动了切片级基础模型(SLFMs)的出现,可在低数据场景下改善预测表现。然而,现有方法未能充分利用任务间与模态间的共享信息。为此,本文提出ModalTune,一种新型微调框架,通过引入模态适配器,在不修改SLFM权重的前提下整合新模态数据。同时,利用大语言模型(LLMs)将标签编码为文本,统一捕捉多个任务及癌种间的语义关系。在四个癌种上,ModalTune相比单模态与多模态模型均达到最先进水平,联合提升生存预测与癌症亚型分类性能,并在泛癌设置下保持竞争力。此外,该方法在两个分布外(OOD)数据集上也展现出良好泛化能力。据我们所知,这是首个面向数字病理中多模态、多任务、泛癌建模的统一微调框架。

原文摘要 · Abstract (English)

Prediction tasks in digital pathology are challenging due to the massive size of whole-slide images (WSIs) and the weak nature of training signals. Advances in computing, data availability, and self-supervised learning (SSL) have paved the way for slide-level foundation models (SLFMs) that can improve prediction tasks in low-data regimes. However, current methods under-utilize shared information between tasks and modalities. To overcome this challenge, we propose ModalTune, a novel fine-tuning framework which introduces the Modal Adapter to integrate new modalities without modifying SLFM weights. Additionally, we use large-language models (LLMs) to encode labels as text, capturing semantic relationships across multiple tasks and cancer types in a single training recipe. ModalTune achieves state-of-the-art (SOTA) results against both uni-modal and multi-modal models across four cancer types, jointly improving survival and cancer subtype prediction while remaining competitive in pan-cancer settings. Additionally, we show ModalTune is generalizable to two out-of-distribution (OOD) datasets. To our knowledge, this is the first unified fine-tuning framework for multi-modal, multi-task, and pan-cancer modeling in digital pathology.

数字病理多模态学习多任务学习基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。