用可学习提示端到端优化病理切片生存分析,提升预测准确率。
End-to-end Multi-source Visual Prompt Tuning for Survival Analysis in Whole Slide Images
- 引入可调视觉提示与适配器,实现生存预测端到端训练
- 在两个免疫组化数据集上C-index分别提升8.7%和12.5%
- 支持多源信息融合,适合病理图像生存分析研究者
基于病理图像的生存分析面临巨大挑战,需从全切片图像(WSIs)中海量图块里定位相关信息。现有方法通常采用两阶段流程:先用预训练网络提取图块特征,再交由生存模型处理。该过程未端到端优化生存模型,且预提取特征未必适合生存预测。为此,本文提出新型端到端视觉提示微调框架VPTSurv,通过高效编码器-解码器结构优化特征嵌入。编码器固定,引入可调视觉提示与适配器,仅优化轻量级适配器和解码器,实现针对生存预测的端到端训练。此外,该框架支持多源信息作为提示,丰富生存模型。VPTSurv在两个免疫组化病理图像数据集上,C-index分别提升8.7%和12.5%,显著优于传统两阶段方法,凸显端到端提示框架的变革潜力。
原文摘要 · Abstract (English)
Survival analysis using pathology images poses a considerable challenge, as it requires the localization of relevant information from the multitude of tiles within whole slide images (WSIs). Current methods typically resort to a two-stage approach, where a pre-trained network extracts features from tiles, which are then used by survival models. This process, however, does not optimize the survival models in an end-to-end manner, and the pre-extracted features may not be ideally suited for survival prediction. To address this limitation, we present a novel end-to-end Visual Prompt Tuning framework for survival analysis, named VPTSurv. VPTSurv refines feature embeddings through an efficient encoder-decoder framework. The encoder remains fixed while the framework introduces tunable visual prompts and adaptors, thus permitting end-to-end training specifically for survival prediction by optimizing only the lightweight adaptors and the decoder. Moreover, the versatile VPTSurv framework accommodates multi-source information as prompts, thereby enriching the survival model. VPTSurv achieves substantial increases of 8.7% and 12.5% in the C-index on two immunohistochemical pathology image datasets. These significant improvements highlight the transformative potential of the end-to-end VPT framework over traditional two-stage methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。