arXiv:2409.17610cs.CLcs.CV2024-09被引 3

用对话上下文提升医疗多轮图文对齐,零样本有效改善模型理解能力。

ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue

  • 利用前序对话内容定位图像兴趣区域,减少干扰信息。
  • 在三个科室测试中显著提升图文对齐效果,统计显著。
  • 适合医疗多模态对话系统研究者与临床AI应用开发者。

近年来大语言模型的快速发展推动了视觉语言模型在医疗领域的广泛应用。在我们的在线医疗咨询场景中,医生需根据患者提供的文本和图像进行多轮对话以诊断病情,形成多轮多模态医疗对话。不同于传统医学视觉问答中由专业设备采集的高质量图像,本场景中的图像由患者手机拍摄,存在背景杂乱、病灶区域严重偏移等问题,导致模型训练阶段视觉-语言对齐性能下降。本文提出ZALM3,一种零样本策略,通过分析图像前的文本对话内容,推断图像中的感兴趣区域(RoIs)。该方法利用大语言模型提取上下文关键词,并结合视觉定位模型提取目标区域,生成去噪后的图像,从而增强视觉-语言对齐。为更精准评估,我们设计了一种新的多轮单模态/多模态医疗对话主观评价指标,实现细粒度性能对比。在三个不同临床科室的实验中,ZALM3表现显著优于基线,具有统计显著性。

原文摘要 · Abstract (English)

The rocketing prosperity of large language models (LLMs) in recent years has boosted the prevalence of vision-language models (VLMs) in the medical sector. In our online medical consultation scenario, a doctor responds to the texts and images provided by a patient in multiple rounds to diagnose her/his health condition, forming a multi-turn multimodal medical dialogue format. Unlike high-quality images captured by professional equipment in traditional medical visual question answering (Med-VQA), the images in our case are taken by patients' mobile phones. These images have poor quality control, with issues such as excessive background elements and the lesion area being significantly off-center, leading to degradation of vision-language alignment in the model training phase. In this paper, we propose ZALM3, a Zero-shot strategy to improve vision-language ALignment in Multi-turn Multimodal Medical dialogue. Since we observe that the preceding text conversations before an image can infer the regions of interest (RoIs) in the image, ZALM3 employs an LLM to summarize the keywords from the preceding context and a visual grounding model to extract the RoIs. The updated images eliminate unnecessary background noise and provide more effective vision-language alignment. To better evaluate our proposed method, we design a new subjective assessment metric for multi-turn unimodal/multimodal medical dialogue to provide a fine-grained performance comparison. Our experiments across three different clinical departments remarkably demonstrate the efficacy of ZALM3 with statistical significance.

医疗AI图文对齐多模态对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。