用小改动伪造临床指令,暴露医学视觉语言模型漏洞。
When Minor Edits Matter: LLM-Driven Prompt Attack for Medical VLM Robustness in Ultrasound
- 用大模型生成贴近真实医嘱的微调指令做攻击
- 主流医学视觉模型在微小提示变化下准确率显著下降
- 适合关注医疗AI安全性的研究人员和开发者
超声因其便携、低成本、安全及实时成像等优势被广泛应用于临床。然而图像采集与解读高度依赖操作者,推动了鲁棒性AI辅助分析方法的发展。视觉语言模型(VLMs)近期在医学影像分析中展现出强大的多模态推理能力,表现接近甚至超越人类。但其可信度面临严峻挑战,尤其是对抗鲁棒性:因医学VLM通过自然语言指令运行,提示词构造成为现实且可利用的脆弱点。微小变化(拼写错误、简写、表述不明确或措辞模糊)可能显著改变模型输出。本文提出一种可扩展的对抗评估框架,利用大语言模型(LLM)通过‘人性化’重写和最小编辑生成临床上合理的对抗提示变体,模仿常规临床交流。基于超声多选题问答基准,系统评估了SOTA医学VLM对这类攻击的脆弱性,分析了攻击者LLM能力对攻击成功率的影响、攻击成功与模型置信度的关系,并识别出跨模型的一致失败模式。结果揭示了亟待解决的真实鲁棒性差距,以确保临床安全落地。代码将在评审后公开。
原文摘要 · Abstract (English)
Ultrasound is widely used in clinical practice due to its portability, cost-effectiveness, safety, and real-time imaging capabilities. However, image acquisition and interpretation remain highly operator dependent, motivating the development of robust AI-assisted analysis methods. Vision-language models (VLMs) have recently demonstrated strong multimodal reasoning capabilities and competitive performance in medical image analysis, including ultrasound. However, emerging evidence highlights significant concerns about their trustworthiness. In particular, adversarial robustness is critical because Med-VLMs operate via natural-language instructions, rendering prompt formulation a realistic and practically exploitable point of vulnerability. Small variations (typos, shorthand, underspecified requests, or ambiguous wording) can meaningfully shift model outputs. We propose a scalable adversarial evaluation framework that leverages a large language model (LLM) to generate clinically plausible adversarial prompt variants via "humanized" rewrites and minimal edits that mimic routine clinical communication. Using ultrasound multiple-choice question answering benchmarks, we systematically assess the vulnerability of SOTA Med-VLMs to these attacks, examine how attacker LLM capacity influences attack success, analyze the relationship between attack success and model confidence, and identify consistent failure patterns across models. Our results highlight realistic robustness gaps that must be addressed for safe clinical translation. Code will be released publicly following the review process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。