arXiv:2603.28508cs.CV2026-03

融合大模型与感知检测,提升AI生成图像识别的泛化能力

Generalizable Detection of AI Generated Images with Large Models and Fuzzy Decision Tree

  • 用模糊决策树融合语义与感知特征,实现多源信息自适应融合
  • 在多个生成模型上达到顶尖准确率,跨模型检测效果显著
  • 适合需要高鲁棒性检测的平台、媒体和内容安全场景

AI生成图像的恶意使用和广泛传播严重威胁数字内容的真实性。现有检测方法依赖生成过程中的低层伪影,但易因模型特异性过拟合而泛化能力差。近期研究尝试使用多模态大语言模型(MLLMs)进行AIGC检测,利用其高层语义推理和广义泛化能力,但缺乏对细微生成痕迹的精细感知,难以独立胜任检测任务。为此,我们提出一种新框架,通过模糊决策树将轻量级感知检测器与MLLMs协同集成。决策树将基础检测器输出视为模糊隶属度,实现语义与感知线索的自适应融合。大量实验表明,该方法在多种生成模型上均达领先精度,并具备强大泛化性能。

原文摘要 · Abstract (English)

The malicious use and widespread dissemination of AI-generated images pose a serious threat to the authenticity of digital content. Existing detection methods exploit low-level artifacts left by common manipulation steps within the generation pipeline, but they often lack generalization due to model-specific overfitting. Recently, researchers have resorted to Multimodal Large Language Models (MLLMs) for AIGC detection, leveraging their high-level semantic reasoning and broad generalization capabilities. While promising, MLLMs lack the fine-grained perceptual sensitivity to subtle generation artifacts, making them inadequate as standalone detectors. To address this issue, we propose a novel AI-generated image detection framework that synergistically integrates lightweight artifact-aware detectors with MLLMs via a fuzzy decision tree. The decision tree treats the outputs of basic detectors as fuzzy membership values, enabling adaptive fusion of complementary cues from semantic and perceptual perspectives. Extensive experiments demonstrate that the proposed method achieves state-of-the-art accuracy and strong generalization across diverse generative models.

AI检测多模态泛化能力模糊逻辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。