医学AI模型对抗攻击评估需多指标,不能只看成功率。
Beyond Attack Success Rate: A Multi-Metric Evaluation of Adversarial Transferability in Medical Imaging Models

- 用多指标评估攻击效果,包括图像质量与扰动强度。
- 发现成功率与图像失真度几乎无关,无法反映真实风险。
- 适合关注医疗AI安全性的研究人员和临床部署团队。
深度学习在医学影像分析中日益普及,但其对对抗扰动的脆弱性引发临床应用担忧。现有评估主要依赖攻击成功率(ASR)这一二元指标,仅判断攻击是否成功,却忽略扰动强度、感知质量及跨架构迁移性等关键因素,导致评估不完整。随着视觉变换器(ViTs)挑战卷积神经网络(CNNs)主导地位,不同架构的学习机制差异使单一指标难以全面刻画对抗行为。为此,我们在四个医学数据集(PathMNIST、DermaMNIST、RetinaMNIST、CheXpert)上,对七种模型(VGG-16、ResNet-50、DenseNet-121、Inception-v3、DeiT、Swin Transformer、ViT-B/16)与七种攻击方法在五种扰动预算下的表现进行了系统评估,测量了ASR、峰值信噪比(PSNR)、结构相似性指数(SSIM)及$L_2$扰动幅度。结果表明,感知质量与失真度高度相关,且与ASR几乎无关,该规律在CNN与ViT中均成立。这说明仅靠ASR无法有效衡量对抗鲁棒性与可迁移性。因此,医疗AI的对抗风险评估必须采用包含攻击效能、方法与代价在内的多指标框架。
原文摘要 · Abstract (English)
While deep learning systems are becoming increasingly prevalent in medical image analysis, their vulnerabilities to adversarial perturbations raise serious concerns for clinical deployment. These vulnerability evaluations largely rely on Attack Success Rate (ASR), a binary metric that indicates solely whether an attack is successful. However, the ASR metric does not account for other factors, such as perturbation strength, perceptual image quality, and cross-architecture attack transferability, and therefore, the interpretation is incomplete. This gap requires consideration, as complex, large-scale deep learning systems, including Vision Transformers (ViTs), are increasingly challenging the dominance of Convolutional Neural Networks (CNNs). These architectures learn differently, and it is unclear whether a single metric, e.g., ASR, can effectively capture adversarial behavior. To address this, we perform a systematic empirical study on four medical image datasets: PathMNIST, DermaMNIST, RetinaMNIST, and CheXpert. We evaluate seven models (VGG-16, ResNet-50, DenseNet-121, Inception-v3, DeiT, Swin Transformer, and ViT-B/16) against seven attack methods at five perturbation budgets, measuring ASR, Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), and $L_2$ perturbation magnitude. Our findings show a consistent pattern: perceptual and distortion metrics are strongly associated with one another and exhibit minimal correlation with ASR. This applies to both CNNs and ViTs. The results demonstrate that ASR alone is an inadequate indicator of adversarial robustness and transferability. Consequently, we argue that a thorough assessment of adversarial risk in medical AI necessitates multi-metric frameworks that encompass not only the attack efficacy but also its methodology and associated overheads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。