arXiv:2504.11922cs.CV2025-04AAAI被引 17

提出新数据集与方法,精准识别局部伪造图像。

Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification Approach

论文配图:Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification Approach
图 1 · 摘自论文原文
  • 构建15万张带场景感知标注的伪造图像数据集
  • 新模型通过噪声指纹放大细微伪造痕迹
  • 适合内容安全、AI检测研究者使用

AI生成图像工具的兴起使得局部伪造愈发逼真,威胁视觉内容真实性。现有数据集多聚焦物体级伪造,忽视天空、地面等大范围区域编辑。为此,我们提出大型数据集BR-Gen,包含15万张具有多样场景感知标注的局部伪造图像,基于语义校准确保样本质量。该数据集通过全自动的“感知-生成-评估”流程构建,保证语义一致性和视觉真实感。同时,提出NFA-ViT模型,一种噪声引导的伪造增强视觉变换器,通过挖掘图像中的异质区域(如潜在编辑区)并利用注意力机制促进正常与异常特征交互,将细微伪造痕迹传播至全图,提升检测鲁棒性。大量实验表明,BR-Gen覆盖了现有方法未涵盖的新场景;NFA-ViT在BR-Gen上表现优于现有方法,并在主流基准上具有良好泛化能力。

原文摘要 · Abstract (English)

The rise of AI-generated image tools has made localized forgeries increasingly realistic, posing challenges for visual content integrity. Although recent efforts have explored localized AIGC detection, existing datasets predominantly focus on object-level forgeries while overlooking broader scene edits in regions such as sky or ground. To address these limitations, we introduce \textbf{BR-Gen}, a large-scale dataset of 150,000 locally forged images with diverse scene-aware annotations, which are based on semantic calibration to ensure high-quality samples. BR-Gen is constructed through a fully automated ``Perception-Creation-Evaluation'' pipeline to ensure semantic coherence and visual realism. In addition, we further propose \textbf{NFA-ViT}, a Noise-guided Forgery Amplification Vision Transformer that enhances the detection of localized forgeries by amplifying subtle forgery-related features across the entire image. NFA-ViT mines heterogeneous regions in images, \emph{i.e.}, potential edited areas, by noise fingerprints. Subsequently, attention mechanism is introduced to compel the interaction between normal and abnormal features, thereby propagating the traces throughout the entire image, allowing subtle forgeries to influence a broader context and improving overall detection robustness. Extensive experiments demonstrate that BR-Gen constructs entirely new scenarios that are not covered by existing methods. Take a step further, NFA-ViT outperforms existing methods on BR-Gen and generalizes well across current benchmarks.

图像伪造检测数据集构建视觉Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。