用SAM模型提升图像伪造定位的跨域泛化能力
SARIF: Segment Anything for Robust Image Forensics

- 基于SAM设计可反馈优化的掩码解码器,自动定位伪造区域
- 在多个基准上实现平均跨数据集性能领先,抗常见图像退化
- 无需人工标注,适合真实场景中的伪造检测应用
图像伪造定位因操作手法多样和分布偏移而困难。现有模型在基准上表现优异,但跨域泛化能力差。本文提出SARIF(Segment Anything for Robust Image Forensics),利用具备提示驱动架构和强泛化能力的SAM。SARIF引入反馈引导的掩码解码器与双编码器设计,提取伪造特异性信息以捕捉取证痕迹,同时保留SAM结构优势。通过块级提示机制,从自适应编码器与冻结编码器的残差特征中提取伪造线索,融合先前掩码提示,驱动反馈式掩码精炼过程,实现无需人工输入的自动伪造分割。在标准伪造定位基准上的大量实验表明,SARIF在跨数据集性能和对常见图像退化鲁棒性方面均表现优异。
原文摘要 · Abstract (English)
Image forgery localization remains challenging due to diverse manipulation techniques and distribution shifts. Existing forgery localization models achieve high accuracy on benchmarks but often struggle with cross-domain generalization and robustness. In this paper, we propose SARIF (Segment Anything for Robust Image Forensics), a framework that leverages the Segment Anything Model (SAM), which has a promptable architecture and strong generalization ability. SARIF introduces a feedback-guided mask decoder and a dual-encoder design that extracts forgery-specific information to capture forensic traces while exploiting the SAM architecture. To localize manipulated regions, we design a block-wise prompting mechanism that derives forgery-specific cues from residual features between an adapted encoder and its frozen counterpart. These features are fused with the previous mask prompt to drive a feedback-based mask refinement process, enabling automatic forgery segmentation without manual input. Extensive experiments on standard forgery-localization benchmarks show that SARIF achieves strong average cross-dataset performance and robustness to common image corruptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。