用文本引导扩散模型生成逼真攻击图,暴露肺部X光诊断模型漏洞。
Text-Guided Diffusion-Based Adversarial Attacks on Chest X-Ray Images

- 通过可学习的文本条件控制扩散生成对抗样本,不直接修改像素。
- 在多病种分类中使模型性能下降至AUROC 0.44-0.49,仍保持高图像质量。
- 医生对攻击图判断无变化,揭示人机诊断差异,适合医学AI安全评估。
随着人工智能在胸部X光(CXR)解读、分诊和临床决策支持中的日益应用,理解其对对抗性操纵的脆弱性对于安全部署至关重要。现有鲁棒性评估主要依赖像素空间攻击,虽有数值限制但未必反映真实的放射学变异。这一局限在多疾病CXR分类中尤为关键,因模型需同时评估多种重叠病理,对抗攻击可能导致多个诊断预测错误。本文提出一种文本引导的扩散对抗框架,优化可学习的文本条件,同时冻结扩散生成器和目标分类器,通过学习到的图像先验生成对抗样本,而非直接像素扰动。我们在多种分类器架构下评估该框架,涵盖二分类肺不张与多疾病分类任务,并与FGSM、PGD及Carlini-Wagner攻击对比。结果表明,该方法始终造成最大性能下降:二分类中AUROC降至0.3885–0.5646,多疾病设置中为0.4441–0.4878;同时保持优异图像保真度(SSIM 0.9080,LPIPS 0.1670,FID 51.23)。重要的是,尽管模型预测发生显著变化,临床医生对95.9%的二分类和73.8%的多疾病攻击图像判断未变。这些发现揭示了人类与机器解读间的临床重要差异,表明医疗AI鲁棒性评估应从传统像素攻击拓展至生成式威胁模型,以暴露在视觉和临床上合理变化下的系统失效。
原文摘要 · Abstract (English)
As artificial intelligence is increasingly integrated into chest X-ray (CXR) interpretation, triage, and clinical decision support, understanding its vulnerability to adversarial manipulation is critical for safe deployment. Existing robustness evaluations, however, predominantly rely on pixel-space attacks that introduce numerically constrained perturbations but may not represent plausible radiographic variation. This limitation is particularly important in multi-disease CXR classification, where models simultaneously evaluate multiple overlapping pathologies and adversarial failures may alter several diagnostic predictions. We propose a text-guided diffusion-based adversarial framework that optimizes learnable text conditioning while keeping the diffusion generator and target classifier frozen, enabling adversarial generation through a learned image prior rather than direct pixel manipulation. We evaluate the framework across multiple classifier architectures in both binary atelectasis and multi-disease CXR classification and compare it with FGSM, PGD, and Carlini-Wagner attacks. Our approach consistently produced the greatest degradation in classifier performance, reducing AUROC to 0.3885-0.5646 in binary classification and 0.4441-0.4878 in the multi-disease setting, while achieving superior image fidelity (SSIM 0.9080, LPIPS 0.1670, FID 51.23). Importantly, clinician interpretation remained unchanged for 95.9% of binary and 73.8% of multi-disease adversarial images despite substantial changes in model predictions. These findings reveal a clinically important discrepancy between human and machine interpretation and demonstrate the need to extend medical AI robustness evaluation beyond conventional pixel-space attacks toward generative threat models that can expose failures under visually and clinically plausible image variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。