用大模型分析假图生成原因,能定位篡改区域并追溯生成方法。
Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics
- 通过精心设计提示词和少量样本学习,激活大模型的伪造分析能力。
- GPT4V在Autosplice和LaMa数据集上准确率分别达92.1%和86.3%。
- 适合图像取证、AI内容检测等需要解释性分析的场景。
生成式AI的快速发展使内容创作更便捷,但也让图像篡改更易实现且难以察觉。尽管多模态大语言模型(LLMs)蕴含丰富世界知识,但其本身并非为对抗人工智能生成内容(AIGC)而设计,难以理解局部篡改细节。本文研究多模态LLMs在伪造检测中的应用,提出一个框架,可评估图像真实性、定位篡改区域、提供证据并根据语义线索追溯生成方法。实验表明,通过精细提示工程与少量样本学习,可有效激发大模型潜力。定性和定量实验显示,GPT4V在Autosplice数据集上准确率达92.1%,在LaMa上达86.3%,表现媲美当前最优AIGC检测方法。文章还讨论了多模态LLMs在此任务中的局限,并提出改进方向。
原文摘要 · Abstract (English)
The rapid development of generative AI facilitates content creation and makes image manipulation easier and more difficult to detect. While multimodal Large Language Models (LLMs) have encoded rich world knowledge, they are not inherently tailored for combating AI-generated Content (AIGC) and struggle to comprehend local forgery details. In this work, we investigate the application of multimodal LLMs in forgery detection. We propose a framework capable of evaluating image authenticity, localizing tampered regions, providing evidence, and tracing generation methods based on semantic tampering clues. Our method demonstrates that the potential of LLMs in forgery analysis can be effectively unlocked through meticulous prompt engineering and the application of few-shot learning techniques. We conduct qualitative and quantitative experiments and show that GPT4V can achieve an accuracy of 92.1% in Autosplice and 86.3% in LaMa, which is competitive with state-of-the-art AIGC detection methods. We further discuss the limitations of multimodal LLMs in such tasks and propose potential improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。