提升医疗视觉问答模型抗攻击能力,同时保持推理可解释性。
SafeMed-R1: Adversarial Reinforcement Learning for Generalizable and Robust Medical Reasoning in Vision-Language Models
- 用对抗训练与分组相对策略优化结合,强化推理过程鲁棒性。
- 在88,000样本上,对抗攻击下仍保持84.45%准确率,比基线高59个百分点。
- 链式思考训练的模型更抗攻击,说明可解释性有助安全防护。
视觉-语言模型(VLMs)在医疗视觉问答(VQA)中展现出巨大潜力,但其在临床部署中受对抗攻击严重威胁。标准对抗训练虽对简单任务有效,却常损害泛化性能和生成的临床推理质量。我们提出SafeMed-R1,一种混合防御框架,在保持高质量、可解释医疗推理的同时确保鲁棒性能。该框架采用两阶段设计:训练时,结合对抗训练与分组相对策略优化(AT-GRPO),显式增强推理过程对最坏情况扰动的鲁棒性;推理时,引入随机平滑,提供$ L_2 $-范数的认证鲁棒性保证。我们在涵盖八种医学影像模态、超过88,000个样本的OmniMedVQA基准上评估SafeMed-R1。实验表明,经微调的标准VLM在干净输入上达95%准确率,但在PGD攻击下骤降至约25%;而SafeMed-R1在相同攻击下仍保持84.45%准确率,鲁棒性提升59个百分点。此外,我们发现显式链式思考训练的模型相比仅指令训练的变体具有更强对抗鲁棒性,表明医疗AI系统中可解释性与安全性存在协同效应。
原文摘要 · Abstract (English)
Vision--Language Models (VLMs) show significant promise for Medical Visual Question Answering (VQA), yet their deployment in clinical settings is hindered by severe vulnerability to adversarial attacks. Standard adversarial training, while effective for simpler tasks, often degrades both generalization performance and the quality of generated clinical reasoning. We introduce SafeMed-R1, a hybrid defense framework that ensures robust performance while preserving high-quality, interpretable medical reasoning. SafeMed-R1 employs a two-stage approach: at training time, we integrate Adversarial Training with Group Relative Policy Optimization (AT-GRPO) to explicitly robustify the reasoning process against worst-case perturbations; at inference time, we augment the model with Randomized Smoothing to provide certified $L_2$-norm robustness guarantees. We evaluate SafeMed-R1 on the OmniMedVQA benchmark across eight medical imaging modalities comprising over 88,000 samples. Our experiments reveal that standard fine-tuned VLMs, despite achieving 95\% accuracy on clean inputs, collapse to approximately 25\% under PGD attacks. In contrast, SafeMed-R1 maintains 84.45\% accuracy under the same adversarial conditions, representing a 59 percentage point improvement in robustness. Furthermore, we demonstrate that models trained with explicit chain-of-thought reasoning exhibit superior adversarial robustness compared to instruction-only variants, suggesting a synergy between interpretability and security in medical AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。