arXiv:2504.07336cs.CVcs.AI2025-04被引 12

用大模型零样本生成医学影像报告,实现多模态分割无需配对数据。

Zeus: Zero-shot LLM Instruction for Union Segmentation in Multimodal Medical Imaging

  • 用冻结的大模型根据影像自动生成诊断级文本指令。
  • 在无需预训练图文数据情况下,多模态分割性能超越主流方法。
  • 适合无标注图文对的医疗场景,尤其适用于放射科临床部署。

医学图像分割因基于UNet和Transformer的骨干网络持续进步而取得显著成果。然而,真实临床诊断常需整合领域知识,尤其是文本信息。多模态学习虽可结合视觉与文本模态,但收集配对的视觉-语言数据成本高、耗时长。受大语言模型(LLM)在跨模态任务中的优异表现启发,我们提出一种新颖的视觉-语言大模型联合框架。具体地,利用冻结的LLM基于对应医学影像(如T1-w或T2-w MRI、CT)生成零样本指令,模拟放射科扫描与报告生成流程。通过提取不同模态的特殊特征并融合信息,实现最终临床诊断支持。借助生成的文本指令,本框架可在无预先收集的视觉-语言数据条件下完成多模态分割。为评估所提方法,我们在多个基准上进行综合实验,统计结果与可视化案例均证明其优越性。

原文摘要 · Abstract (English)

Medical image segmentation has achieved remarkable success through the continuous advancement of UNet-based and Transformer-based foundation backbones. However, clinical diagnosis in the real world often requires integrating domain knowledge, especially textual information. Conducting multimodal learning involves visual and text modalities shown as a solution, but collecting paired vision-language datasets is expensive and time-consuming, posing significant challenges. Inspired by the superior ability in numerous cross-modal tasks for Large Language Models (LLMs), we proposed a novel Vision-LLM union framework to address the issues. Specifically, we introduce frozen LLMs for zero-shot instruction generation based on corresponding medical images, imitating the radiology scanning and report generation process. {To better approximate real-world diagnostic processes}, we generate more precise text instruction from multimodal radiology images (e.g., T1-w or T2-w MRI and CT). Based on the impressive ability of semantic understanding and rich knowledge of LLMs. This process emphasizes extracting special features from different modalities and reunion the information for the ultimate clinical diagnostic. With generated text instruction, our proposed union segmentation framework can handle multimodal segmentation without prior collected vision-language datasets. To evaluate our proposed method, we conduct comprehensive experiments with influential baselines, the statistical results and the visualized case study demonstrate the superiority of our novel method.}

多模态分割零样本医学影像大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。