arXiv:2605.26421cs.CV2026-05被引 1

动态调整提示词,更好识别伪造图像。

HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image Detection

论文配图:HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image Detection
图 1 · 摘自论文原文
  • 真实图像用统一提示锚定,假图像用自适应提示捕捉差异。
  • 在多个基准上达到当前最优检测效果。
  • 适合需要精准区分不同伪造手法的场景。

生成模型的快速发展导致虚假内容泛滥,给现有合成图像检测(SID)方法带来挑战。尽管近期工作利用视觉语言模型(如CLIP)和可学习文本提示来识别合成图像,但这些方法仍使用静态提示作为真实与虚假图像的固定边界,无法适应推理时出现的不同伪造类型。为此,我们提出HydraPrompt:一种非对称提示框架,通过与细粒度图像线索对齐,动态调整类别中心。具体地,我们设计了非对称提示适配器(APA):(1) 对真实类别,引入一组统一提示以捕捉一致的代表性模式,作为真实内容的统一锚点;(2) 对虚假类别,构建样本自适应提示,专门捕捉不同样本中的多样化线索,实现对伪造图像变异的自适应建模。为进一步增强不同合成图像间的判别力,我们引入条件监督对比(CSC)目标,压缩真实表示并捕获细粒度伪造特征。在多个主流SID基准上的实验表明,该框架性能达到当前最优水平。

原文摘要 · Abstract (English)

The rapid evolution of generative models has precipitated a proliferation of fabricated content, posing significant challenges to existing Synthetic Image Detection (SID) methods. Capitalizing on advancements in vision-language models (e.g., CLIP), recent attempts have leveraged learnable textual prompts to identify synthetic images. However, they still leverage static prompt as a fixed boundary for real and fake images, failing to adapt to the varying types of forgery that emerge during inference. To overcome this issue, we propose **HydraPrompt**, an asymmetric prompting framework that dynamically adjusts the category centers by aligning with fine-grained image cues. Specifically, we propose an Asymmetric Prompt Adapter (**APA**): (1) for authentic category, we introduce a single set of prompts to capture the consistent representative patterns, which serves as a unified anchor for real content. While (2) for fake category, we construct sample-adaptive prompts that specialize in capturing diverse cues from different samples, enabling adaptive modeling of forgery image variations. To increase pronounced discriminability within different synthetic images, we further introduce a Conditional Supervised Contrastive (**CSC**) objective, which compacts the authentic representations while capturing fine-grained forgery clues. Extensive experiments on popular SID benchmarks demonstrate the state-of-the-art performance of our framework.

图像检测视觉语言模型伪造识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。