提出不丢像素的高分辨率伪造图像检测框架,提升真实场景识别能力。
No Pixel Left Behind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image Detection
- 分块融合全分辨率局部特征与全局视图,确保细节不丢失。
- 在Chameleon数据集上准确率提升超13%,HiRes-50K上提升10%。
- 适合关注高分辨伪造检测、真实性分析的研究者使用。
高分辨率、精细生成的AI图像快速增长,对现有检测方法构成严峻挑战。当前方法多在低分辨率自动数据集上训练评估,常通过缩放或中心裁剪适配输入,但会遗漏部分像素,导致细微高频伪影或信息丢失。本文提出高分辨率细节聚合网络(HiDA-Net),采用特征聚合模块(FAM)融合多个全分辨率局部块与下采样全局视图特征,并与全局表示融合进行最终预测,实现原生分辨率细节的保留与利用。为增强对局部篡改和压缩的鲁棒性,引入逐标记伪造定位(TFL)模块以提升空间敏感度,以及JPEG质量因子估计(QFE)模块显式分离生成伪影与压缩噪声。此外,构建新基准HiRes-50K,包含50,568张最高达64兆像素的图像。大量实验表明,HiDA-Net在挑战性数据集Chameleon上准确率提升超过13%,在自建的HiRes-50K上提升10%。
原文摘要 · Abstract (English)
The rapid growth of high-resolution, meticulously crafted AI-generated images poses a significant challenge to existing detection methods, which are often trained and evaluated on low-resolution, automatically generated datasets that do not align with the complexities of high-resolution scenarios. A common practice is to resize or center-crop high-resolution images to fit standard network inputs. However, without full coverage of all pixels, such strategies risk either obscuring subtle, high-frequency artifacts or discarding information from uncovered regions, leading to input information loss. In this paper, we introduce the High-Resolution Detail-Aggregation Network (HiDA-Net), a novel framework that ensures no pixel is left behind. We use the Feature Aggregation Module (FAM), which fuses features from multiple full-resolution local tiles with a down-sampled global view of the image. These local features are aggregated and fused with global representations for final prediction, ensuring that native-resolution details are preserved and utilized for detection. To enhance robustness against challenges such as localized AI manipulations and compression, we introduce Token-wise Forgery Localization (TFL) module for fine-grained spatial sensitivity and JPEG Quality Factor Estimation (QFE) module to disentangle generative artifacts from compression noise explicitly. Furthermore, to facilitate future research, we introduce HiRes-50K, a new challenging benchmark consisting of 50,568 images with up to 64 megapixels. Extensive experiments show that HiDA-Net achieves state-of-the-art, increasing accuracy by over 13% on the challenging Chameleon dataset and 10% on our HiRes-50K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。