发现跨生成范式图像伪造新线索,提升检测泛化能力
Ghosts Beneath Textures: Texture-Relation Cues for Cross-Paradigm AI-Generated Image Detection

- 通过抑制语义干扰,首次揭示跨范式共有的纹理结构关系
- 提出DTS-Det框架,99.6%准确率,较最优基线提升10.5%
- 适用于图像条件与非条件生成场景,抗重构与对抗攻击
AI生成图像迅速泛滥,现有检测器多针对噪声或文本引导等无图生成范式设计,但图像条件生成在实际应用中日益重要。本文构建了首个跨范式检测基准ConImageGen,发现现有方法在两类生成范式间泛化失败。为此,研究首次可视化了语义无关的纹理模式,其呈现局部-全局结构化关系,构成通用伪造证据。基于此,提出DTS-Det框架,聚焦纹理关系建模而非显式伪影。大量实验验证:DTS-Det在ConImageGen上达99.6%准确率,较最佳基线提升10.5%;在PicoBanana/RAID跨数据集测试中分别达93.2%/94.1%;面对重构攻击与黑盒对抗攻击,检测率仍保持95.2%/88.1%。
原文摘要 · Abstract (English)
AI-generated images have proliferated rapidly, motivating extensive research. Most existing AI-generated image detectors are developed and evaluated under image-free generation paradigms, such as noise-based or text-guided generation. However, image-conditioned generation has become increasingly important in practical applications, as it enables more fine-grained control over generated content. Detecting AI-generated images across these two paradigms creates a critical cross-paradigm detection problem that has long been overlooked. To study this problem, we construct ConImageGen, a benchmark for cross-paradigm AI-generated image detection. Evaluations on ConImageGen show that existing detectors fail to generalize reliably across image-free and image-conditioned generation. To address this failure, this paper identifies a cross-paradigm forensic cue and provides a new perspective for generalized AI-generated image detection. Specifically, by suppressing semantic interference, we visualize, for the first time, semantics-irrelevant texture patterns across generation paradigms. These patterns exhibit structured local-global texture relations, indicating a generalizable form of forensic evidence. Motivated by this finding, we shift the focus from directly exploiting explicit artifacts to modeling texture relations and propose DTS-Det, a detection framework that captures and leverages such relations for generalized AI-generated image detection. Extensive experiments validate the effectiveness of our method. DTS-Det achieves state-of-the-art performance across diverse evaluation settings, reaching 99.6% ACC on ConImageGen with a 10.5% gain over the best baseline. It also achieves 93.2%/94.1% ACC in cross-dataset evaluation on PicoBanana/RAID and maintains detection rates of 95.2%/88.1% under reconstruction attacks and black-box adversarial attacks, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。