基于CT影像与文本的肾癌多任务模型,提升诊断与预后预测精度。
A Disease-Centric Vision-Language Foundation Model for Precision Oncology in Kidney Cancer
- 采用两阶段预训练,融合医学知识增强图像与文本表征
- 在10项任务中超越现有模型,生存预测C-index达0.726(提升约20%)
- 仅需20%数据即可达到最优性能,适合临床少样本场景
非侵入性评估日益增多的偶然发现性肾肿块是泌尿肿瘤学中的关键挑战,诊断不确定性常导致良性或惰性肿瘤的过度治疗。本研究构建并验证了RenalCLIP,该模型基于来自9家中国医疗中心和公开TCIA队列的27,866例CT扫描(8,809名患者),是一个用于肾肿块表征、诊断与预后的视觉-语言基础模型。模型通过两阶段预训练策略,先用领域知识增强图像与文本编码器,再通过对比学习对齐,生成具有强泛化能力的鲁棒表征。RenalCLIP在涵盖肾癌全临床流程的10个核心任务中表现更优,包括解剖评估、诊断分类与生存预测,尤其在TCIA队列的复发无病生存预测任务中,C-index达0.726,较领先基线提升约20%。此外,其预训练带来显著数据效率:在诊断分类任务中,仅需20%训练数据即可达到所有基线模型在全量数据微调后的峰值性能。同时,在报告生成、图文检索与零样本诊断任务中也表现优异。结果表明,RenalCLIP为提升肾癌诊断准确性、优化预后分层及个性化管理提供了有力工具。
原文摘要 · Abstract (English)
The non-invasive assessment of increasingly incidentally discovered renal masses is a critical challenge in urologic oncology, where diagnostic uncertainty frequently leads to the overtreatment of benign or indolent tumors. In this study, we developed and validated RenalCLIP using a dataset of 27,866 CT scans from 8,809 patients across nine Chinese medical centers and the public TCIA cohort, a visual-language foundation model for characterization, diagnosis and prognosis of renal mass. The model was developed via a two-stage pre-training strategy that first enhances the image and text encoders with domain-specific knowledge before aligning them through a contrastive learning objective, to create robust representations for superior generalization and diagnostic precision. RenalCLIP achieved better performance and superior generalizability across 10 core tasks spanning the full clinical workflow of kidney cancer, including anatomical assessment, diagnostic classification, and survival prediction, compared with other state-of-the-art general-purpose CT foundation models. Especially, for complicated task like recurrence-free survival prediction in the TCIA cohort, RenalCLIP achieved a C-index of 0.726, representing a substantial improvement of approximately 20% over the leading baselines. Furthermore, RenalCLIP's pre-training imparted remarkable data efficiency; in the diagnostic classification task, it only needs 20% training data to achieve the peak performance of all baseline models even after they were fully fine-tuned on 100% of the data. Additionally, it achieved superior performance in report generation, image-text retrieval and zero-shot diagnosis tasks. Our findings establish that RenalCLIP provides a robust tool with the potential to enhance diagnostic accuracy, refine prognostic stratification, and personalize the management of patients with kidney cancer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。