首个面向医学影像的区域描述框架,能精准生成病灶部位的临床级描述。
Describe Anything in Medical Images
- 用医学专家设计的提示词引导大模型进行区域描述。
- 在3个数据集上超越GPT-4o等顶尖模型,临床事实性更优。
- 适合临床辅助诊断与医学视觉语言模型研究者使用。
局部图像描述技术已取得显著进展,如描述任何模型(DAM)可在无显式区域-文本监督下生成区域特异性描述。然而,此类能力尚未广泛应用于医学影像领域,而该领域诊断依赖于细微区域发现而非全局理解。为此,我们提出MedDAM,首个基于大视觉-语言模型的医学图像区域描述综合框架。MedDAM采用医学专家设计的模态特定提示词,并建立包含定制评估协议、数据预处理流程和专用问答模板库的稳健评估基准。该基准通过属性级验证任务评估MedDAM及其他可适配的大视觉-语言模型的临床事实性,克服了医学数据集中缺乏真实区域-描述对的问题。在VinDr-CXR、LIDC-IDRI和SkinCon数据集上的大量实验表明,MedDAM优于包括GPT-4o、Claude 3.7 Sonnet、LLaMA-3.2 Vision、Qwen2.5-VL、GPT-4Rol和OMG-LLaVA在内的领先模型,揭示了区域语义对齐在医学图像理解中的重要性,并确立了MedDAM作为临床视觉-语言融合的有前景基础。
原文摘要 · Abstract (English)
Localized image captioning has made significant progress with models like the Describe Anything Model (DAM), which can generate detailed region-specific descriptions without explicit region-text supervision. However, such capabilities have yet to be widely applied to specialized domains like medical imaging, where diagnostic interpretation relies on subtle regional findings rather than global understanding. To mitigate this gap, we propose MedDAM, the first comprehensive framework leveraging large vision-language models for region-specific captioning in medical images. MedDAM employs medical expert-designed prompts tailored to specific imaging modalities and establishes a robust evaluation benchmark comprising a customized assessment protocol, data pre-processing pipeline, and specialized QA template library. This benchmark evaluates both MedDAM and other adaptable large vision-language models, focusing on clinical factuality through attribute-level verification tasks, thereby circumventing the absence of ground-truth region-caption pairs in medical datasets. Extensive experiments on the VinDr-CXR, LIDC-IDRI, and SkinCon datasets demonstrate MedDAM's superiority over leading peers (including GPT-4o, Claude 3.7 Sonnet, LLaMA-3.2 Vision, Qwen2.5-VL, GPT-4Rol, and OMG-LLaVA) in the task, revealing the importance of region-level semantic alignment in medical image understanding and establishing MedDAM as a promising foundation for clinical vision-language integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。