arXiv:2511.06947cs.CVcs.AI2025-11

提出伪造CLIP评分的特征空间错位方法,可让图像骗过质量评估但保持视觉清晰。

FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection

  • 通过特征对齐与分布平衡,生成能误导CLIP评分的对抗图像。
  • 在10个艺术提示和ImageNet子集上,CLIPscore显著提升且视觉保真度高。
  • 发现灰度化导致特征退化,据此提出91%准确率的篡改检测新方法。

基于CLIP的模型具备良好的图文对齐特性,使其广泛应用于CLIPscore等图像质量评估。然而,这种对齐机制极易被破坏。本文提出FoCLIP,一种用于欺骗基于CLIP的图像质量评估的特征空间错位框架。基于随机梯度下降,FoCLIP整合三个核心组件:以特征对齐为核心模块减少图文模态差异,分数分布平衡模块与像素保护正则化共同优化CLIPscore表现与图像质量间的多模态均衡。该设计可使生成图像在多种输入提示下最大化CLIPscore,即使从人类感知角度看存在视觉不可识别或语义不符。在10个艺术杰作提示及ImageNet子集上的实验表明,优化图像在显著提升CLIPscore的同时保持高视觉保真度。此外,我们发现灰度转换会引发伪造图像特征显著退化,导致CLIPscore明显下降,但统计特性仍与原图一致。受此启发,我们提出一种基于色彩通道敏感性的篡改检测机制,在标准基准上达到91%准确率。本工作为基于CLIP的多模态系统中的特征错位现象及其防御提供了实用路径。

原文摘要 · Abstract (English)

The well-aligned attribute of CLIP-based models enables its effective application like CLIPscore as a widely adopted image quality assessment metric. However, such a CLIP-based metric is vulnerable for its delicate multimodal alignment. In this work, we propose \textbf{FoCLIP}, a feature-space misalignment framework for fooling CLIP-based image quality metric. Based on the stochastic gradient descent technique, FoCLIP integrates three key components to construct fooling examples: feature alignment as the core module to reduce image-text modality gaps, the score distribution balance module and pixel-guard regularization, which collectively optimize multimodal output equilibrium between CLIPscore performance and image quality. Such a design can be engineered to maximize the CLIPscore predictions across diverse input prompts, despite exhibiting either visual unrecognizability or semantic incongruence with the corresponding adversarial prompts from human perceptual perspectives. Experiments on ten artistic masterpiece prompts and ImageNet subsets demonstrate that optimized images can achieve significant improvement in CLIPscore while preserving high visual fidelity. In addition, we found that grayscale conversion induces significant feature degradation in fooling images, exhibiting noticeable CLIPscore reduction while preserving statistical consistency with original images. Inspired by this phenomenon, we propose a color channel sensitivity-driven tampering detection mechanism that achieves 91% accuracy on standard benchmarks. In conclusion, this work establishes a practical pathway for feature misalignment in CLIP-based multimodal systems and the corresponding defense method.

图像伪造CLIP对抗攻击检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。