为医疗视觉语言模型设计安全评分框架,可自动评估漏洞风险。
VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models
- 构建攻击模板与不可察觉的视觉扰动,模拟真实临床威胁。
- 通过双LLM评分生成0-32分风险综合指标,平均漏洞提升超0.6σ。
- 开源工具支持一键测试,适合医疗AI安全研究者使用。
视觉语言模型(VLMs)有望简化耗时的医学影像工作流程,但其在临床环境中的系统性安全评估仍十分匮乏。本文提出VSF-Med,一个端到端的医疗VLM漏洞评分框架,包含三个创新组件:(i) 针对新兴威胁向量的丰富文本提示攻击模板库;(ii) 基于结构相似性(SSIM)阈值校准的不可察觉视觉扰动,以保持临床真实性;(iii) 由两名独立判断的LLM评估的八维评分体系,原始分数经z-score归一化后生成0–32的综合风险评分。该框架完全基于公开数据集构建,并提供开源代码,从5,000张放射科图像中合成超过30,000个对抗样本,支持任何医疗VLM的单命令可复现基准测试。综合分析显示,主流VLM在持续攻击效应、提示注入有效性及安全绕过成功率上平均提升分别为0.90σ、0.74σ、0.63σ。其中,Llama-3.2-11B-Vision-Instruct在持续攻击效应上峰值提升达1.29σ,而GPT-4o在同类向量上提升0.69σ,提示注入攻击提升0.28σ。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) hold great promise for streamlining labour-intensive medical imaging workflows, yet systematic security evaluations in clinical settings remain scarce. We introduce VSF--Med, an end-to-end vulnerability-scoring framework for medical VLMs that unites three novel components: (i) a rich library of sophisticated text-prompt attack templates targeting emerging threat vectors; (ii) imperceptible visual perturbations calibrated by structural similarity (SSIM) thresholds to preserve clinical realism; and (iii) an eight-dimensional rubric evaluated by two independent judge LLMs, whose raw scores are consolidated via z-score normalization to yield a 0--32 composite risk metric. Built entirely on publicly available datasets and accompanied by open-source code, VSF--Med synthesizes over 30,000 adversarial variants from 5,000 radiology images and enables reproducible benchmarking of any medical VLM with a single command. Our consolidated analysis reports mean z-score shifts of $0.90σ$ for persistence-of-attack-effects, $0.74σ$ for prompt-injection effectiveness, and $0.63σ$ for safety-bypass success across state-of-the-art VLMs. Notably, Llama-3.2-11B-Vision-Instruct exhibits a peak vulnerability increase of $1.29σ$ for persistence-of-attack-effects, while GPT-4o shows increases of $0.69σ$ for that same vector and $0.28σ$ for prompt-injection attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。