arXiv:2501.02385cs.CVcs.CL2025-01中稿 · NAACL被引 10

用视觉提示引导医学图文模型关注特定区域,提升诊断准确性。

Guiding Medical Vision-Language Models with Explicit Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations

  • 用简单视觉标记指导模型聚焦医学图像特定区域。
  • 在多个医疗问答数据集上超越现有顶尖模型表现。
  • 适合临床辅助诊断系统开发与医学AI可解释性研究者。

主流视觉语言模型虽能理解图像整体信息,但难以根据人类指定区域进行聚焦。现有方法依赖大量高质量图文配对数据来学习注意力分布。为此,本文提出利用视觉提示——多种形式的简单视觉标记,引导并增强模型生成特定区域的注意力。我们设计了首个医学视觉提示框架MedVP,整合医学实体识别、视觉提示生成与数据集适配,实现提示引导的微调。实验表明,该方法在多个医疗视觉问答数据集上显著优于当前最优模型。通过大规模实验与人工评估,分析了不同提示形式对性能的影响,验证了方法的有效性与临床价值。

原文摘要 · Abstract (English)

While mainstream vision-language models (VLMs) have advanced rapidly in understanding image level information, they still lack the ability to focus on specific areas designated by humans. Rather, they typically rely on large volumes of high-quality image-text paired data to learn and generate posterior attention maps. To address this critical issue, we propose leveraging visual prompts:simple visual markers in various forms to guide and enhance the formation of region-specific attention. Thus, we introduce MedVP, a pioneering framework that integrates medical entity extraction, visual prompt generation, and dataset adaptation for visual prompt guided fine-tuning. We successfully outperform recent state-of-the-art large models across multiple medical VQA datasets. Extensive experiments and Human evaluation are conducted to analyze the impact of different visual prompt forms and how they contribute to performance improvement. The results demonstrate both the effectiveness and clinical significance of our approach.

医学AI视觉提示多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。