arXiv:2604.27875cs.CV2026-04

提出FGINet模型,提升AI生成图像检测的泛化能力。

Frequency-Aware Semantic Fusion with Gated Injection for AI-generated Image Detection

论文配图:Frequency-Aware Semantic Fusion with Gated Injection for AI-generated Image Detection
图 1 · 摘自论文原文
  • 用频域掩码减少对特定生成器的依赖,增强泛化性。
  • 分层门控注入频率特征,缓解语义与频率表示冲突。
  • 适合需要跨模型检测生成图像的研究者使用。

AI生成图像日益逼真多样,给通用检测带来挑战。尽管视觉基础模型(VFMs)提供丰富语义信息,基于频率的方法捕捉互补的伪影线索,但现有融合方法在未见生成模型上性能显著下降。我们归因于两个关键因素:频率捷径偏差(过度依赖特定生成器的易区分线索)和高层语义与低层频率模式间的跨域表示冲突。为此,我们提出频率感知门控注入网络(FGINet)。设计频带掩码编码器(BMFE),在频域实施跨频带掩码,降低对生成器特异性模式的依赖,促进更多样、泛化的表示。引入分层门控频率注入(LGFI)机制,以自适应门控方式逐步将频率线索注入VFM主干,匹配其层次抽象特性,缓解表示冲突。此外,提出超球紧凑性学习(HCL)框架,采用余弦间隔目标,学习紧凑且分离良好的表示。大量实验表明,FGINet在多个挑战性数据集上达到领先性能并具备强泛化能力。

原文摘要 · Abstract (English)

AI-generated images are becoming increasingly realistic and diverse, posing significant challenges for generalizable detection. While Vision Foundation Models (VFMs) provide rich semantic representations and frequency-based methods capture complementary artifact cues, existing approaches that combine these modalities still suffer from limited generalization, with notable performance degradation on unseen generative models. We attribute this limitation to two key factors: frequency shortcut bias toward easily distinguishable cues associated with specific generators and cross-domain representation conflict between high-level semantics and low-level frequency patterns. To address these issues, we propose a Frequency-aware Gated Injection Network (FGINet) to improve generalization. Specifically, we design a Band-Masked Frequency Encoder (BMFE) that applies cross-band masking in the frequency domain to reduce reliance on generator-specific patterns and encourage more diverse and generalizable representations. We further introduce a Layer-wise Gated Frequency Injection (LGFI) mechanism to progressively inject frequency cues into the VFM backbone with adaptive gating, aligning with its hierarchical abstraction and alleviating representation conflict. Moreover, we propose a Hyperspherical Compactness Learning (HCL) framework with a cosine margin objective to learn compact and well-separated representations. Extensive experiments demonstrate that FGINet achieves state-of-the-art performance and strong generalization across multiple challenging datasets.

图像检测生成模型频域分析泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。