arXiv:2509.10250cs.CV2025-09被引 1

通过多任务与篡改增强训练,提升图像生成检测的泛化能力

GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection

  • 引入多种篡改策略保持内容一致性
  • 多任务监督实现跨域像素级溯源,准确率提升5.8%
  • 适合需要跨模型检测生成图像的研究者

随着生成模型日益复杂多样,检测AI生成图像面临更大挑战。现有检测器在分布内生成图像上表现良好,但对未见过的生成模型泛化能力有限,主要因其依赖特定生成痕迹(如风格先验、压缩模式)。为此,我们提出GAMMA框架,通过多样化篡改策略(如基于修复的篡改、语义保持扰动)确保篡改前后内容一致,并采用多任务监督机制,包含双分割头与分类头,实现跨生成域的像素级来源归因。此外,引入反向交叉注意力机制,使分割头能引导并修正分类分支中的偏差表示。该方法在GenImage基准上达到当前最优泛化性能,准确率提升5.8%,同时对新发布的生成模型(如GPT-4o)也表现出强鲁棒性。

原文摘要 · Abstract (English)

With generative models becoming increasingly sophisticated and diverse, detecting AI-generated images has become increasingly challenging. While existing AI-genereted Image detectors achieve promising performance on in-distribution generated images, their generalization to unseen generative models remains limited. This limitation is largely attributed to their reliance on generation-specific artifacts, such as stylistic priors and compression patterns. To address these limitations, we propose GAMMA, a novel training framework designed to reduce domain bias and enhance semantic alignment. GAMMA introduces diverse manipulation strategies, such as inpainting-based manipulation and semantics-preserving perturbations, to ensure consistency between manipulated and authentic content. We employ multi-task supervision with dual segmentation heads and a classification head, enabling pixel-level source attribution across diverse generative domains. In addition, a reverse cross-attention mechanism is introduced to allow the segmentation heads to guide and correct biased representations in the classification branch. Our method achieves state-of-the-art generalization performance on the GenImage benchmark, imporving accuracy by 5.8%, but also maintains strong robustness on newly released generative model such as GPT-4o.

图像检测生成模型泛化能力多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。