arXiv:2605.14799cs.CVcs.CR2026-05

测试视觉Mamba在识别AI生成图像上的表现,发现它有潜力但仍有局限。

Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation

论文配图:Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation
图 1 · 摘自论文原文
  • 用视觉Mamba模型对比CNN、ViT和VLM,系统评估其检测能力。
  • 在多个数据集上验证,其准确率与效率优于部分传统模型。
  • 适合关注AI内容安全、需高效检测的开发者和研究者参考。

近年来,计算机视觉因卷积神经网络(CNN)、生成对抗网络(GAN)、基于扩散的架构、视觉变换器(ViTs)以及近期的视觉语言模型(VLMs)等创新架构而取得显著进展,推动了高度逼真且多样化的视觉内容生成。然而,这些生成技术也带来了误用风险,如虚假信息传播、身份盗用及隐私安全威胁。与此同时,基于Mamba的架构已在图像分类、分割、医学影像、目标检测和图像修复等任务中展现多功能性,但在识别AI生成图像方面的潜力仍远未被充分探索。本研究对多种视觉Mamba变体进行了系统性评估,与代表性CNN、ViT和基于VLM的检测器在多样数据集和合成图像源上进行对比分析,重点考察准确率、效率及跨图像类型和生成模型的泛化能力。结果揭示了视觉Mamba在真实性检测中的优势与当前局限,对构建可信赖的视觉内容鉴别系统具有重要意义。

原文摘要 · Abstract (English)

In recent years, computer vision has witnessed remarkable progress, fueled by the development of innovative architectures such as Convolutional Neural Networks (CNNs), Generative Adversarial Networks (GANs), diffusion-based architectures, Vision Transformers (ViTs), and, more recently, Vision-Language Models (VLMs). This progress has undeniably contributed to creating increasingly realistic and diverse visual content. However, such advancements in image generation also raise concerns about potential misuse in areas such as misinformation, identity theft, and threats to privacy and security. In parallel, Mamba-based architectures have emerged as versatile tools for a range of image analysis tasks, including classification, segmentation, medical imaging, object detection, and image restoration, in this rapidly evolving field. However, their potential for identifying AI-generated images remains relatively unexplored compared to established techniques. This study provides a systematic evaluation and comparative analysis of Vision Mamba models for AI-generated image detection. We benchmark multiple Vision Mamba variants against representative CNNs, ViTs, and VLM-based detectors across diverse datasets and synthetic image sources, focusing on key metrics such as accuracy, efficiency, and generalizability across diverse image types and generative models. Through this comprehensive analysis, we aim to elucidate Vision Mamba's strengths and limitations relative to established methodologies in terms of applicability, accuracy, and efficiency in detecting AI-generated images. Overall, our findings highlight both the promise and current limitations of Vision Mamba as a component in systems designed to distinguish authentic from AI-generated visual content. This research is crucial for enhancing detection in an age where distinguishing between real and AI-generated content is a major challenge.

AI检测视觉Mamba图像生成内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。