提出全局伪造先验GAP-SAM,提升AI图像篡改定位泛化能力
GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation Localization

- 用冻结VAE重建图生成全局伪造特征,注入SAM特征金字塔
- 跨6个数据集平均像素F1达79.8,领先现有方法12.6点
- 对压缩、模糊、缩放等常见退化鲁棒性强,适合实际应用
AI生成图像篡改定位需识别被编辑像素,但其跨域性能低于图像级检测,原因在于像素级监督使取证证据与数据集特有的掩码几何和语义边界纠缠。为提升定位泛化性,本文构建基于源图轮廓与深度图的COCO-ControlNet,对齐语义与几何,改善多个定位器的跨域表现。然而,更紧密的掩码-变分自编码器重建对齐(Mask-VAE)表现更差,表明VAE重构伪影难以迁移到扩散修复伪影。我们还发现‘边界黏附’现象:微调的分割模型将预测强制贴合语义对象轮廓而非真实篡改边界。据此提出GAP-SAM,将图像及其冻结的VAE重建图编码为全局伪造令牌,通过零门控FiLM注入SAM3的特征金字塔,在像素解码前进行调节。该机制不依赖空间区域指定,能抑制语义边界捷径,保留定位精度。在六个数据集上,GAP-SAM平均像素F1达79.8,优于最强基线12.6点,且在不同压缩、高斯模糊、重采样程度下均表现最优。
原文摘要 · Abstract (English)
AI-generated image manipulation localization identifies edited pixels, but its OOD performance lags behind image-level detection partly because pixel supervision entangles forensic evidence with dataset-specific mask geometry and semantic boundaries. Extending image-level distribution alignment to localization, we construct COCO-ControlNet with source-image Canny edges and depth maps to align semantics and geometry, improving OOD performance across multiple localizers. Yet tighter Mask-VAE Reconstruction Alignment (Mask-VAE) underperforms COCO-ControlNet, showing that VAE reconstruction artifacts transfer poorly to local diffusion-inpainting artifacts. We also identify \emph{boundary adhesion}, where fine-tuned segmentation models snap predictions to semantic object contours rather than true manipulation boundaries. These findings motivate GAP-SAM, which encodes an image and its frozen VAE reconstruction into a global artifact token and injects it into SAM3's feature pyramid via zero-gated FiLM before pixel decoding. Without prescribing a spatial region, this token modulates dense decoding to preserve localization while suppressing semantic-boundary shortcuts. Across six datasets, GAP-SAM averages 79.8 Pixel-F1, outperforming the strongest prior method by 12.6 points. It also performs best at every tested severity of JPEG compression, Gaussian blur, and resizing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。