arXiv:2409.00543cs.CVcs.CL2024-09中稿 · NeurIPS被引 2

研究不同提示对医疗视觉语言模型零样本任务的影响

How Does Diverse Interpretability of Textual Prompts Impact Medical Vision-Language Zero-Shot Tasks?

  • 设计六种临床真实提示风格评估模型表现
  • 三种模型在不同提示下性能波动大,稳定性差
  • 提示可解释性越高,模型越难理解复杂医学概念

近期医疗视觉语言预训练(MedVLP)在图像分类等零样本任务中取得显著进展,主要依赖大规模医图像-文本对预训练。然而,类别描述的文本提示多样性会显著影响任务性能,亟需模型具备对多样提示风格的鲁棒性。但这一敏感性尚未被充分研究。本文首次系统评估三种主流MedVLP方法在15种疾病上的提示敏感性,设计六种反映真实临床场景的提示风格,并按可解释性排序。结果表明,所有评估模型在不同提示风格下表现不稳定,显示其鲁棒性不足;且随着提示可解释性提升,模型性能反而下降,反映出理解复杂医学概念的困难。该研究强调需进一步改进MedVLP方法以增强对多样化零样本提示的适应能力。

原文摘要 · Abstract (English)

Recent advancements in medical vision-language pre-training (MedVLP) have significantly enhanced zero-shot medical vision tasks such as image classification by leveraging large-scale medical image-text pair pre-training. However, the performance of these tasks can be heavily influenced by the variability in textual prompts describing the categories, necessitating robustness in MedVLP models to diverse prompt styles. Yet, this sensitivity remains underexplored. In this work, we are the first to systematically assess the sensitivity of three widely-used MedVLP methods to a variety of prompts across 15 different diseases. To achieve this, we designed six unique prompt styles to mirror real clinical scenarios, which were subsequently ranked by interpretability. Our findings indicate that all MedVLP models evaluated show unstable performance across different prompt styles, suggesting a lack of robustness. Additionally, the models' performance varied with increasing prompt interpretability, revealing difficulties in comprehending complex medical concepts. This study underscores the need for further development in MedVLP methodologies to enhance their robustness to diverse zero-shot prompts.

医疗AI零样本学习提示工程视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。