一种通用且可迁移的攻击,让病理模型识别失效
Universal and Transferable Attacks on Pathology Foundation Models
- 用固定噪声模式干扰多种病理模型的特征表示
- 在不同视野和数据集上均导致模型性能显著下降
- 适合关注AI安全与防御的研究者阅读
我们提出针对病理基础模型的通用可迁移对抗扰动(UTAP),揭示其关键脆弱性。通过深度学习优化,UTAP采用固定且微弱的噪声模式,添加至病理图像后,系统性破坏多个病理基础模型的特征表示能力,导致下游任务性能下降,包括对广泛未见数据分布的误分类。此外,我们验证了UTAP的两大特性:(1) 通用性——扰动可跨不同视场应用,不受训练数据集限制;(2) 可迁移性——能有效降低从未见过的外部黑盒病理基础模型性能。这两项特性表明,UTAP并非特定于某模型或数据集,而是对多种新兴病理基础模型及其应用构成普遍威胁。我们在多个数据集上系统评估了多种先进病理基础模型,仅以肉眼不可察觉的输入修改,即造成显著性能下降。该攻击的开发为模型鲁棒性评估建立了高标准基准,凸显提升防御机制的必要性,并可能为对抗训练提供关键资源,以保障人工智能在病理学中的安全可靠部署。
原文摘要 · Abstract (English)
We introduce Universal and Transferable Adversarial Perturbations (UTAP) for pathology foundation models that reveal critical vulnerabilities in their capabilities. Optimized using deep learning, UTAP comprises a fixed and weak noise pattern that, when added to a pathology image, systematically disrupts the feature representation capabilities of multiple pathology foundation models. Therefore, UTAP induces performance drops in downstream tasks that utilize foundation models, including misclassification across a wide range of unseen data distributions. In addition to compromising the model performance, we demonstrate two key features of UTAP: (1) universality: its perturbation can be applied across diverse field-of-views independent of the dataset that UTAP was developed on, and (2) transferability: its perturbation can successfully degrade the performance of various external, black-box pathology foundation models - never seen before. These two features indicate that UTAP is not a dedicated attack associated with a specific foundation model or image dataset, but rather constitutes a broad threat to various emerging pathology foundation models and their applications. We systematically evaluated UTAP across various state-of-the-art pathology foundation models on multiple datasets, causing a significant drop in their performance with visually imperceptible modifications to the input images using a fixed noise pattern. The development of these potent attacks establishes a critical, high-standard benchmark for model robustness evaluation, highlighting a need for advancing defense mechanisms and potentially providing the necessary assets for adversarial training to ensure the safe and reliable deployment of AI in pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。