提出一种抗多重干扰补丁的视觉变压器,提升医疗等关键场景的模型可靠性。
Filtered-ViT: A Robust Defense Against Multiple Adversarial Patch Attacks
- 引入自适应多尺度滤波机制,智能识别并抑制受损区域。
- 在四块1%面积补丁攻击下,鲁棒准确率达46.3%,优于现有方法。
- 首次实现对对抗性与自然伪影的统一防御,适合医疗影像等高危应用。
深度学习视觉系统在医疗等安全关键领域日益广泛应用,但仍易受小型对抗补丁的影响而误分类。现有防御大多假设仅存在单个补丁,当多个局部干扰同时出现时失效,而这正是攻击者和真实世界噪声常利用的情况。我们提出Filtered-ViT,一种集成SMART向量中值滤波(SMART-VMF)的新型视觉变换器架构,该机制具有空间自适应、多尺度、鲁棒感知特性,可在保留语义细节的同时选择性抑制被污染区域。在ImageNet上面对LaVAN多补丁攻击,Filtered-ViT在四块1%面积补丁同时存在时,保持79.8%干净准确率和46.3%鲁棒准确率,显著优于现有防御方案。此外,在放射科医学图像的真实案例研究中,Filtered-ViT有效缓解了遮挡、扫描噪声等自然伪影,且未损害诊断内容。这标志着Filtered-ViT是首个在对抗性与自然补丁类干扰上均展现统一鲁棒性的视觉变换器,为真正高风险环境中的可靠视觉系统提供了新路径。
原文摘要 · Abstract (English)
Deep learning vision systems are increasingly deployed in safety-critical domains such as healthcare, yet they remain vulnerable to small adversarial patches that can trigger misclassifications. Most existing defenses assume a single patch and fail when multiple localized disruptions occur, the type of scenario adversaries and real-world artifacts often exploit. We propose Filtered-ViT, a new vision transformer architecture that integrates SMART Vector Median Filtering (SMART-VMF), a spatially adaptive, multi-scale, robustness-aware mechanism that enables selective suppression of corrupted regions while preserving semantic detail. On ImageNet with LaVAN multi-patch attacks, Filtered-ViT achieves 79.8% clean accuracy and 46.3% robust accuracy under four simultaneous 1\% patches, outperforming existing defenses. Beyond synthetic benchmarks, a real-world case study on radiographic medical imagery shows that Filtered-ViT mitigates natural artifacts such as occlusions and scanner noise without degrading diagnostic content. This establishes Filtered-ViT as the first transformer to demonstrate unified robustness against both adversarial and naturally occurring patch-like disruptions, charting a path toward reliable vision systems in truly high-stakes environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。