arXiv:2604.18250cs.CV2026-04

用视觉指令微调提升CT影像生存预测,让模型懂医图还能说人话。

Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning

论文配图:Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning
图 1 · 摘自论文原文
  • 通过医学影像与报告配对,用指令微调训练多模态模型
  • 在临床数据弱预测时,生存预测准确率显著优于基线方法
  • 既可做预测,又能生成医生级的图文解释,适合临床辅助

精准预后判断和风险评估对指导临床决策和优化患者管理至关重要。尽管放射科医生从CT扫描中提取的特征能有效反映疾病严重程度和预后,但图像解读需专业知识,将丰富视觉信息转化为文本摘要不可避免导致信息丢失。本文提出一种基于3D CT图像理解的视觉-语言框架,利用大规模开源CT图像与放射科报告对,通过视觉指令微调进行预训练。该预训练使模型学习到具有临床意义的视觉-文本表征,进而可适配至下游生存预测任务。在预训练模型基础上添加生存预测头,本方法在结合CT影像与临床数据时,提升了生存预测性能,并能对预设问题生成具有临床意义的语言响应。实验表明,该方法在临床数据预测能力较弱时表现尤为突出,优于各类基线方法。代码将在论文接受后发布。

原文摘要 · Abstract (English)

Accurate prognostication and risk estimation are essential for guiding clinical decision-making and optimizing patient management. While radiologist-assessed features from CT scans provide valuable indicators of disease severity and outcomes, interpreting such images requires expert knowledge, and translating rich visual information into textual summaries inevitably leads to information loss. In this work, we propose a vision-language framework for 3D CT image understanding that leverages large-scale open-sourced CT images paired with radiology reports through visual instruction tuning. This pre-training enables the model to learn clinically meaningful visual-textual representations, which can then be adapted to downstream survival prediction tasks. By incorporating a survival prediction head on top of the pre-trained model, our approach improves survival prediction from CT images and clinical data while generating clinically meaningful language responses to predefined questions. Experimental results demonstrate that our method outperforms baseline methods in survival prediction, particularly, when clinical data alone is less predictive. The code will be released upon acceptance.

医学影像生存预测视觉语言指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。