用大模型检测假图并解释造假痕迹,让机器判断更透明。
Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation

- 基于多模态大模型,直接识别合成图像并生成自然语言解释。
- 在超10万张图像上验证,分类与解释能力均达领先水平。
- 适合需要可解释性的内容审核、媒体真实性验证场景。
随着人工智能生成内容(AIGC)技术的快速发展,合成图像日益普及,对真实性的评估和检测带来新挑战。现有方法虽能有效判断图像真伪并定位伪造区域,但普遍缺乏人类可理解的解释,难以应对日益复杂的合成数据。为此,我们提出 FakeVLM——一种专用于通用合成图像与 DeepFake 检测的大规模多模态模型。FakeVLM 不仅能准确区分真实与虚假图像,还能提供清晰、自然语言描述的图像伪造特征说明,显著提升可解释性。同时,我们构建了 FakeClue 数据集,包含超过 10 万张图像,覆盖七个类别,每张图像均配有细粒度的自然语言伪造线索标注。实验表明,FakeVLM 在多个数据集上的表现媲美专家级模型,且无需额外分类器,是合成数据检测的稳健方案。大规模评估证实其在真实性分类与伪造特征解释任务中均具优势,树立了新的基准。代码、模型权重与数据集详见:https://github.com/opendatalab/FakeVLM。
原文摘要 · Abstract (English)
With the rapid advancement of Artificial Intelligence Generated Content (AIGC) technologies, synthetic images have become increasingly prevalent in everyday life, posing new challenges for authenticity assessment and detection. Despite the effectiveness of existing methods in evaluating image authenticity and locating forgeries, these approaches often lack human interpretability and do not fully address the growing complexity of synthetic data. To tackle these challenges, we introduce FakeVLM, a specialized large multimodal model designed for both general synthetic image and DeepFake detection tasks. FakeVLM not only excels in distinguishing real from fake images but also provides clear, natural language explanations for image artifacts, enhancing interpretability. Additionally, we present FakeClue, a comprehensive dataset containing over 100,000 images across seven categories, annotated with fine-grained artifact clues in natural language. FakeVLM demonstrates performance comparable to expert models while eliminating the need for additional classifiers, making it a robust solution for synthetic data detection. Extensive evaluations across multiple datasets confirm the superiority of FakeVLM in both authenticity classification and artifact explanation tasks, setting a new benchmark for synthetic image detection. The code, model weights, and dataset can be found here: https://github.com/opendatalab/FakeVLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。