提出FASA框架,统一定位传统与扩散生成的图像篡改
Bridging the Micro--Macro Gap: Frequency-Aware Semantic Alignment for Image Manipulation Localization
- 用自适应双频DCT提取篡改敏感频域特征
- 在冻结CLIP上做补丁级对比对齐,学习语义先验
- 适合检测扩散模型生成的局部真实篡改,抗退化
随着生成式图像编辑的发展,图像篡改定位(IML)需同时应对传统篡改的明显取证痕迹和扩散模型生成编辑的局部真实感。现有方法通常仅依赖低层取证线索或高层语义,导致微观与宏观特征间的根本性差距。为此,我们提出FASA,一种统一定位传统与扩散生成篡改的框架。具体而言,通过自适应双频DCT模块提取篡改敏感频率特征,并在冻结的CLIP表示上进行补丁级对比对齐,学习篡改感知的语义先验。随后,通过语义-频率侧适配器将这些先验注入分层频率路径,实现多尺度特征交互;再使用原型引导、频率门控的掩码解码器,融合语义一致性与边界感知,完成篡改区域预测。在OpenSDI及多个传统篡改基准上的大量实验表明,该方法达到最先进定位性能,具备强跨生成器与跨数据集泛化能力,且在常见图像退化下表现稳健。
原文摘要 · Abstract (English)
As generative image editing advances, image manipulation localization (IML) must handle both traditional manipulations with conspicuous forensic artifacts and diffusion-generated edits that appear locally realistic. Existing methods typically rely on either low-level forensic cues or high-level semantics alone, leading to a fundamental micro--macro gap. To bridge this gap, we propose FASA, a unified framework for localizing both traditional and diffusion-generated manipulations. Specifically, we extract manipulation-sensitive frequency cues through an adaptive dual-band DCT module and learn manipulation-aware semantic priors via patch-level contrastive alignment on frozen CLIP representations. We then inject these priors into a hierarchical frequency pathway through a semantic-frequency side adapter for multi-scale feature interaction, and employ a prototype-guided, frequency-gated mask decoder to integrate semantic consistency with boundary-aware localization for tampered region prediction. Extensive experiments on OpenSDI and multiple traditional manipulation benchmarks demonstrate state-of-the-art localization performance, strong cross-generator and cross-dataset generalization, and robust performance under common image degradations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。