用病理报告辅助增强病理切片生存分析,提升癌症预后预测效果。
Enhancing WSI-Based Survival Analysis with Report-Auxiliary Self-Distillation
- 利用大模型从报告中提取与切片相关文本,指导特征筛选。
- 通过自蒸馏过滤无关切片特征,在TCGA-BRCA上提高风险预测准确率。
- 适合关注医学图像与文本融合、肿瘤预后研究的科研人员。
基于全切片图像(WSIs)的生存分析对评估癌症预后至关重要,因其提供详尽的微观信息。然而,传统方法常面临特征噪声和数据可及性受限的问题,难以有效捕捉关键预后特征。尽管病理报告包含丰富的患者特异性信息,但其在增强WSI生存分析中的潜力尚未被充分挖掘。为此,本文提出一种新型报告辅助自蒸馏框架(Rasa)。首先,借助先进大语言模型(LLMs),通过精心设计的任务提示从原始嘈杂的病理报告中提取细粒度、与切片相关的文本描述;其次,设计基于自蒸馏的流程,在教师模型的文本知识引导下,过滤学生模型中的无关或冗余切片特征;最后,在学生模型训练中引入风险感知混合策略,提升训练数据的数量与多样性。在自建数据集(CRC)和公开数据集(TCGA-BRCA)上的大量实验表明,Rasa显著优于当前主流方法。代码已开源。
原文摘要 · Abstract (English)
Survival analysis based on Whole Slide Images (WSIs) is crucial for evaluating cancer prognosis, as they offer detailed microscopic information essential for predicting patient outcomes. However, traditional WSI-based survival analysis usually faces noisy features and limited data accessibility, hindering their ability to capture critical prognostic features effectively. Although pathology reports provide rich patient-specific information that could assist analysis, their potential to enhance WSI-based survival analysis remains largely unexplored. To this end, this paper proposes a novel Report-auxiliary self-distillation (Rasa) framework for WSI-based survival analysis. First, advanced large language models (LLMs) are utilized to extract fine-grained, WSI-relevant textual descriptions from original noisy pathology reports via a carefully designed task prompt. Next, a self-distillation-based pipeline is designed to filter out irrelevant or redundant WSI features for the student model under the guidance of the teacher model's textual knowledge. Finally, a risk-aware mix-up strategy is incorporated during the training of the student model to enhance both the quantity and diversity of the training data. Extensive experiments carried out on our collected data (CRC) and public data (TCGA-BRCA) demonstrate the superior effectiveness of Rasa against state-of-the-art methods. Our code is available at https://github.com/zhengwang9/Rasa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。