arXiv:2506.05358cs.CVcs.AI2025-06被引 1

GPT-4V无需微调即可识别图像拼接篡改,表现接近专业模型。

Can ChatGPT Perform Image Splicing Detection? A Preliminary Study

  • 用零样本、少样本和思维链提示评估GPT-4V的图像拼接检测能力。
  • 零样本下准确率超85%,思维链提示在真假图像间平衡最佳。
  • 能结合视觉细节与现实常识,适合做可解释的图像取证辅助工具。

多模态大语言模型(如GPT-4V)具备跨文本与图像模态推理的能力,在复杂视觉语言任务中展现潜力。本初步研究考察GPT-4V在图像伪造检测领域,特别是图像拼接篡改检测中的开箱即用能力。在未进行任何任务特定微调的情况下,我们采用三种提示策略(零样本、少样本、思维链)在经过筛选的CASIA v2.0拼接数据集子集上评估GPT-4V。结果表明,其在零样本设置下达到超过85%的准确率,其中思维链提示在真实与篡改图像间表现出最均衡的性能。定性分析进一步显示,该模型不仅能检测低层视觉伪影,还能利用真实世界上下文知识,如物体尺度、语义一致性及建筑常识,识别不合理的合成图像。尽管GPT-4V仍落后于专门的顶尖拼接检测模型,但其泛化能力、可解释性以及百科式推理能力,凸显其在图像取证中作为灵活工具的潜力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) like GPT-4V are capable of reasoning across text and image modalities, showing promise in a variety of complex vision-language tasks. In this preliminary study, we investigate the out-of-the-box capabilities of GPT-4V in the domain of image forensics, specifically, in detecting image splicing manipulations. Without any task-specific fine-tuning, we evaluate GPT-4V using three prompting strategies: Zero-Shot (ZS), Few-Shot (FS), and Chain-of-Thought (CoT), applied over a curated subset of the CASIA v2.0 splicing dataset. Our results show that GPT-4V achieves competitive detection performance in zero-shot settings (more than 85% accuracy), with CoT prompting yielding the most balanced trade-off across authentic and spliced images. Qualitative analysis further reveals that the model not only detects low-level visual artifacts but also draws upon real-world contextual knowledge such as object scale, semantic consistency, and architectural facts, to identify implausible composites. While GPT-4V lags behind specialized state-of-the-art splicing detection models, its generalizability, interpretability, and encyclopedic reasoning highlight its potential as a flexible tool in image forensics.

图像伪造大模型视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。