arXiv:2510.23775cs.CVcs.AI2025-10

用轻量模型和视觉语言模型实现边缘设备上图像真伪可解释检测

Explainable Detection of AI-Generated Images with Artifact Localization Using Faster-Than-Lies and Vision-Language Models for Edge Devices

  • 结合轻量卷积分类器与视觉语言模型进行联合判断
  • 在32×32图像上达96.5%准确率,推理仅175ms
  • 生成热图定位缺陷并自动生成解释文本

AI生成图像的真实性日益难以辨别。本文提出一种可解释的图像真伪检测系统,结合轻量级卷积分类器(Faster-Than-Lies)与视觉语言模型(Qwen2-VL-7B),对32×32图像进行分类、定位和解释。模型在扩展的CiFAKE数据集(含对抗扰动)上达到96.5%准确率,且在8核CPU上推理时间仅为175ms,适用于本地或边缘设备部署。通过自编码器重建误差图生成缺陷定位热图,提升人机可解释性。我们将70种视觉伪影归类为8个语义类别,并实现对每类异常的可解释文本生成。该工作验证了视觉与语言推理融合在低分辨率图像真伪检测中的可行性,拓展了其在取证、工业检测及社交媒体审核等领域的应用潜力。

原文摘要 · Abstract (English)

The increasing realism of AI-generated imagery poses challenges for verifying visual authenticity. We present an explainable image authenticity detection system that combines a lightweight convolutional classifier ("Faster-Than-Lies") with a Vision-Language Model (Qwen2-VL-7B) to classify, localize, and explain artifacts in 32x32 images. Our model achieves 96.5% accuracy on the extended CiFAKE dataset augmented with adversarial perturbations and maintains an inference time of 175ms on 8-core CPUs, enabling deployment on local or edge devices. Using autoencoder-based reconstruction error maps, we generate artifact localization heatmaps, which enhance interpretability for both humans and the VLM. We further categorize 70 visual artifact types into eight semantic groups and demonstrate explainable text generation for each detected anomaly. This work highlights the feasibility of combining visual and linguistic reasoning for interpretable authenticity detection in low-resolution imagery and outlines potential cross-domain applications in forensics, industrial inspection, and social media moderation.

图像检测可解释性边缘计算视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。