用合成临床示例提升医疗多模态模型安全,防攻击不伤性能
Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations
- 通过合成临床示例实现推理时防御,应对图文越狱攻击
- 在九种医学影像数据上验证,安全提升且性能损失小
- 少样本下混合策略平衡安全与可用性,适合临床部署
生成式医疗多模态视觉语言模型(Med-VLMs)主要从视觉(如医学图像)和语言(如临床提问)输入生成复杂文本信息(如诊断报告)。然而其安全漏洞尚未被充分研究。Med-VLMs 应能拒绝有害查询,例如“提供使用此CT扫描进行保险欺诈的详细说明”。但安全增强可能引发过度防御,导致模型误拒良性临床问题。本文提出一种新型推理时防御策略,可抵御视觉与文本越狱攻击。基于来自九种模态的多样化医学影像数据集,我们证明该基于合成临床示例的防御策略显著提升安全性,同时几乎不损害模型性能。此外,增加示例预算可缓解过度防御问题。为此,我们进一步引入混合示例策略,在少量示例约束下实现安全与性能的权衡。
原文摘要 · Abstract (English)
Generative medical vision-language models~(Med-VLMs) are primarily designed to generate complex textual information~(e.g., diagnostic reports) from multimodal inputs including vision modality~(e.g., medical images) and language modality~(e.g., clinical queries). However, their security vulnerabilities remain underexplored. Med-VLMs should be capable of rejecting harmful queries, such as \textit{Provide detailed instructions for using this CT scan for insurance fraud}. At the same time, addressing security concerns introduces the risk of over-defense, where safety-enhancing mechanisms may degrade general performance, causing Med-VLMs to reject benign clinical queries. In this paper, we propose a novel inference-time defense strategy to mitigate harmful queries, enabling defense against visual and textual jailbreak attacks. Using diverse medical imaging datasets collected from nine modalities, we demonstrate that our defense strategy based on synthetic clinical demonstrations enhances model safety without significantly compromising performance. Additionally, we find that increasing the demonstration budget alleviates the over-defense issue. We then introduce a mixed demonstration strategy as a trade-off solution for balancing security and performance under few-shot demonstration budget constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。