arXiv:2502.11725cs.CVcs.LG2025-02中稿 · publication in the…被引 14

让CLIP模型更抗攻击,能生成更可靠的视觉相似度度量

Adversarially Robust CLIP Models Can Induce Better (Robust) Perceptual Metrics

  • 通过无监督对抗微调构建抗干扰的R-CLIP$_\textrm{F}$模型
  • 在零样本设置下超越现有度量,且对扰动保持高鲁棒性
  • 适合需高可靠性的图像检索与NSFW内容过滤场景

衡量视觉相似性是计算机视觉的关键工具。近年来,基于大型多样化数据集训练的神经网络(如CLIP)提取特征的感知度量受到青睐。然而,这类度量不具备对抗鲁棒性。本文展示,通过无监督对抗微调得到的抗干扰CLIP模型R-CLIP$_\textrm{F}$,可诱导出更优且具备鲁棒性的感知度量,在零样本设置下性能优于现有方法,并在微调后达到当前最优水平。此外,该度量在鲁棒图像到图像检索任务中表现优异,尤其适用于“不适合工作场所”(NSFW)内容检测与数据集过滤。标准度量易受微小扰动攻击导致NSFW检测失效,而本方法在攻击下仍保持高准确率,且未受扰动图像性能相当。同时,鲁棒CLIP模型诱导的度量具有更高可解释性:特征反演可揭示哪些图像被视作相似,文本反演可找到与特定提示关联的图像,从而可视化模型学习到的丰富视觉概念,包括记忆的人脸、画作及复杂查询。

原文摘要 · Abstract (English)

Measuring perceptual similarity is a key tool in computer vision. In recent years perceptual metrics based on features extracted from neural networks with large and diverse training sets, e.g. CLIP, have become popular. At the same time, the metrics extracted from features of neural networks are not adversarially robust. In this paper we show that adversarially robust CLIP models, called R-CLIP$_\textrm{F}$, obtained by unsupervised adversarial fine-tuning induce a better and adversarially robust perceptual metric that outperforms existing metrics in a zero-shot setting, and further matches the performance of state-of-the-art metrics while being robust after fine-tuning. Moreover, our perceptual metric achieves strong performance on related tasks such as robust image-to-image retrieval, which becomes especially relevant when applied to "Not Safe for Work" (NSFW) content detection and dataset filtering. While standard perceptual metrics can be easily attacked by a small perturbation completely degrading NSFW detection, our robust perceptual metric maintains high accuracy under an attack while having similar performance for unperturbed images. Finally, perceptual metrics induced by robust CLIP models have higher interpretability: feature inversion can show which images are considered similar, while text inversion can find what images are associated to a given prompt. This also allows us to visualize the very rich visual concepts learned by a CLIP model, including memorized persons, paintings and complex queries.

感知度量对抗鲁棒CLIP可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。