arXiv:2510.03483cs.CVcs.AI2025-10

DuPLUS用双提示机制实现跨模态医学图像分割与预后预测,通用性强且易适配新任务。

DuPLUS: Dual-Prompt Vision-Language Framework for Universal Medical Image Segmentation and Prognosis

  • 设计双提示框架,通过分层文本控制实现精细任务调节。
  • 在10个数据集上分割表现超越主流模型,头颈癌预后预测CI达0.69。
  • 支持快速适配新任务和医院数据,适合临床部署与多场景应用。

深度学习在医学影像分析中受限于任务专用模型的泛化能力不足及缺乏预后预测能力,而现有‘通用’方法则存在条件设定简单、医学语义理解薄弱的问题。为此,我们提出DuPLUS,一种高效多模态医学图像分析的视觉-语言框架。该框架引入分层语义提示机制,实现对分析任务的细粒度控制,这是以往通用模型所缺失的能力。为提升可扩展性,其采用由独特双提示机制驱动的分层文本控制架构。在分割任务中,杜普拉斯可在三种成像模态、十个解剖结构各异的数据集上泛化,覆盖30余种器官与肿瘤类型,且在10个数据集中的8个上优于当前最优的任务专用与通用模型。我们进一步展示其文本控制架构的可扩展性,无缝集成电子健康记录(EHR)数据用于预后预测,在头颈部癌症数据集上达到0.69的协和指数(CI)。参数高效的微调使其能快速适应不同中心的新任务与模态,确立了杜普拉斯作为通用且临床相关的医学图像分析解决方案。代码已公开:https://anonymous.4open.science/r/DuPLUS-6C52

原文摘要 · Abstract (English)

Deep learning for medical imaging is hampered by task-specific models that lack generalizability and prognostic capabilities, while existing 'universal' approaches suffer from simplistic conditioning and poor medical semantic understanding. To address these limitations, we introduce DuPLUS, a deep learning framework for efficient multi-modal medical image analysis. DuPLUS introduces a novel vision-language framework that leverages hierarchical semantic prompts for fine-grained control over the analysis task, a capability absent in prior universal models. To enable extensibility to other medical tasks, it includes a hierarchical, text-controlled architecture driven by a unique dual-prompt mechanism. For segmentation, DuPLUS is able to generalize across three imaging modalities, ten different anatomically various medical datasets, encompassing more than 30 organs and tumor types. It outperforms the state-of-the-art task specific and universal models on 8 out of 10 datasets. We demonstrate extensibility of its text-controlled architecture by seamless integration of electronic health record (EHR) data for prognosis prediction, and on a head and neck cancer dataset, DuPLUS achieved a Concordance Index (CI) of 0.69. Parameter-efficient fine-tuning enables rapid adaptation to new tasks and modalities from varying centers, establishing DuPLUS as a versatile and clinically relevant solution for medical image analysis. The code for this work is made available at: https://anonymous.4open.science/r/DuPLUS-6C52

医学影像双提示分割预后预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。