让伪造检测模型的判断与解释自洽出错,暴露其脆弱性。
JECA^2: Judgment-Explanation Consistent Adversarial Attack against Forensic Vision-Language Models

- 用梯度引导扰动转移视觉注意力至正常区域
- 优化文本提示嵌入使解释与目标判断一致
- 可诊断模型一致性缺陷,适合安全评估者使用
伪造视觉语言模型(VLMs)近年被用于检测图像篡改并提供自然语言解释,但其对抗鲁棒性仍待深入研究。现有攻击多仅改变模型二元判断,伴随解释仍可能暴露伪造线索且与判断矛盾。本文提出JECA^2,一种可控白盒红队诊断方法,联合调整视觉归因与文本解释,使其与目标判断保持一致。在视觉端,采用Grad-CAM引导的扰动将归因从篡改区域转移到正常区域;在文本端,在词元邻近约束下优化提示嵌入,使其趋向真实性的语义表达。在多个伪造VLM基准测试中,JECA^2在白盒设置下实现了更高的攻击成功率和自动化判断-解释一致性,对闭源模型的迁移攻击虽存在但有限。结果揭示了基于解释的伪造检测模型存在一致性失效模式,呼吁未来评估超越二元检测准确率。
原文摘要 · Abstract (English)
Forensic vision-language models (VLMs) have recently been developed to detect image tampering and provide natural-language explanations. However, their robustness against adversarial manipulation remains underexplored. Existing adversarial attacks typically aim to flip the model's binary judgment, while the accompanying explanation may still reveal forensic cues and contradict the attacked judgment. In this paper, we study judgment-explanation consistent adversarial attacks against forensic VLMs and propose JECA^2, a controlled white-box red-team diagnostic that jointly redirects visual attribution and aligns textual explanations with the target judgment. On the visual side, JECA^2 uses Grad-CAM-guided perturbations to divert attribution from tampered regions toward benign regions. On the textual side, it optimizes prompt embeddings toward authenticity-affirming semantics under a token-proximity constraint. Experiments on forensic VLM benchmarks show that JECA^2 achieves higher attack success and automated judgment-explanation consistency than implemented baselines under white-box threat settings, while transfer to closed-source VLMs remains measurable but limited. Our results highlight a consistency failure mode in explanation-based forensic VLMs and motivate future robustness evaluation beyond binary detection accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。