通过分层视觉语言协作,提升癌症生存预测精度。
HiLa: Hierarchical Vision-Language Collaboration for Cancer Survival Prediction
- 分层提取病理切片的局部与区域特征,结合多维度语言提示增强对齐。
- 在三个TCGA数据集上达到当前最优性能,显著提升预测准确率。
- 适合医学图像分析、生存预测及多模态学习研究者参考。
基于全切片图像(WSIs)的生存预测在癌症研究中至关重要。尽管已有显著进展,现有方法受限于稀疏的切片级标签,难以从千兆像素级的WSI中学习判别性表征。近年来,融合语言监督的视觉语言(VL)模型展现出潜力,但其在生存预测中的应用仍不充分。主要原因有二:一是现有方法仅依赖单一语言提示和基础余弦相似度,无法捕捉多维度语言信息与视觉特征之间的细粒度关联,导致视觉-语言对齐不足;二是多数方法仅利用补丁级信息,忽视了WSI的固有层级结构及其层级间交互,难以有效建模层次化关系。为此,我们提出一种新型分层视觉语言协作框架(HiLa)。HiLa采用预训练特征提取器,在补丁和区域两个层级生成分层视觉特征,并在每一层级构造描述不同生存相关属性的语言提示,通过最优提示学习(OPL)实现与视觉特征的对齐。该方法可全面学习对应不同生存属性的判别性视觉特征,从而改善视觉-语言对齐。此外,引入跨层级传播(CLP)与互对比学习(MCL)模块,促进补丁与区域层级间的交互与一致性,最大化层级协作。在三个TCGA数据集上的实验表明,该方法性能达到当前最优水平。
原文摘要 · Abstract (English)
Survival prediction using whole-slide images (WSIs) is crucial in cancer re-search. Despite notable success, existing approaches are limited by their reliance on sparse slide-level labels, which hinders the learning of discriminative repre-sentations from gigapixel WSIs. Recently, vision language (VL) models, which incorporate additional language supervision, have emerged as a promising solu-tion. However, VL-based survival prediction remains largely unexplored due to two key challenges. First, current methods often rely on only one simple lan-guage prompt and basic cosine similarity, which fails to learn fine-grained associ-ations between multi-faceted linguistic information and visual features within WSI, resulting in inadequate vision-language alignment. Second, these methods primarily exploit patch-level information, overlooking the intrinsic hierarchy of WSIs and their interactions, causing ineffective modeling of hierarchical interac-tions. To tackle these problems, we propose a novel Hierarchical vision-Language collaboration (HiLa) framework for improved survival prediction. Specifically, HiLa employs pretrained feature extractors to generate hierarchical visual features from WSIs at both patch and region levels. At each level, a series of language prompts describing various survival-related attributes are constructed and aligned with visual features via Optimal Prompt Learning (OPL). This ap-proach enables the comprehensive learning of discriminative visual features cor-responding to different survival-related attributes from prompts, thereby improv-ing vision-language alignment. Furthermore, we introduce two modules, i.e., Cross-Level Propagation (CLP) and Mutual Contrastive Learning (MCL) to maximize hierarchical cooperation by promoting interactions and consistency be-tween patch and region levels. Experiments on three TCGA datasets demonstrate our SOTA performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。