给医疗多模态大模型加安全防护,不降性能还防越狱攻击
The Forgotten Shield: Safety Grafting in Parameter-Space for Medical MLLMs
- 在参数空间内注入原始模型的安全知识,实现高效重对齐
- 实验显示安全防护显著提升,医疗任务性能下降小于5%
- 适合关注医疗AI安全、部署落地的研究者与工程师
医学多模态大模型(Medical MLLMs)在专业医疗任务中取得显著进展,但其安全性研究滞后,存在真实部署风险。本文首次构建多维度评估框架,系统评测当前SOTA Medical MLLMs的安全性。实证分析发现,现有模型在通用与医学特定安全维度均普遍存在漏洞,尤其易受跨模态越狱攻击。此外,医学微调常导致模型原有安全对齐的灾难性遗忘。为此,我们提出一种新型“参数空间干预”方法,从原始基模型中提取内在安全知识,并在构建医学能力过程中同步注入目标模型。同时设计细粒度参数搜索算法,在安全与医疗性能间实现最优平衡。实验表明,该方法无需额外领域安全数据即可显著增强医学大模型的安全防护,核心医疗性能损失低于5%。
原文摘要 · Abstract (English)
Medical Multimodal Large Language Models (Medical MLLMs) have achieved remarkable progress in specialized medical tasks; however, research into their safety has lagged, posing potential risks for real-world deployment. In this paper, we first establish a multidimensional evaluation framework to systematically benchmark the safety of current SOTA Medical MLLMs. Our empirical analysis reveals pervasive vulnerabilities across both general and medical-specific safety dimensions in existing models, particularly highlighting their fragility against cross-modality jailbreak attacks. Furthermore, we find that the medical fine-tuning process frequently induces catastrophic forgetting of the model's original safety alignment. To address this challenge, we propose a novel "Parameter-Space Intervention" approach for efficient safety re-alignment. This method extracts intrinsic safety knowledge representations from original base models and concurrently injects them into the target model during the construction of medical capabilities. Additionally, we design a fine-grained parameter search algorithm to achieve an optimal trade-off between safety and medical performance. Experimental results demonstrate that our approach significantly bolsters the safety guardrails of Medical MLLMs without relying on additional domain-specific safety data, while minimizing degradation to core medical performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。