用现成视觉模型做图像篡改定位,无需重训练。
Off-the-shelf Vision Models Benefit Image Manipulation Localization
- 用可训练适配器复用现成视觉模型的语义先验。
- 在多个数据集上准确率超现有方法,最高提升3.2%。
- 适合想快速部署篡改定位系统的研究者与工程师。
图像篡改定位(IML)与通用视觉任务通常被视为两个独立方向,因其操纵特定特征与语义特征存在根本差异。本文提出新视角:二者本质关联,通用语义先验可助力IML。基于此,我们设计了一种新型可训练适配器(ReVi),将现有现成通用视觉模型(如图像生成与分割网络)用于IML。受鲁棒主成分分析启发,该适配器从模型中解耦出语义冗余信息,仅增强操纵相关特征。与需大量重设计和全量训练的现有方法不同,本方法冻结原模型参数,仅微调适配器。实验表明,该方法性能优越,在多个数据集上表现领先,最高提升3.2%,展现出可扩展的IML框架潜力。
原文摘要 · Abstract (English)
Image manipulation localization (IML) and general vision tasks are typically treated as two separate research directions due to the fundamental differences between manipulation-specific and semantic features. In this paper, however, we bridge this gap by introducing a fresh perspective: these two directions are intrinsically connected, and general semantic priors can benefit IML. Building on this insight, we propose a novel trainable adapter (named ReVi) that repurposes existing off-the-shelf general-purpose vision models (e.g., image generation and segmentation networks) for IML. Inspired by robust principal component analysis, the adapter disentangles semantic redundancy from manipulation-specific information embedded in these models and selectively enhances the latter. Unlike existing IML methods that require extensive model redesign and full retraining, our method relies on the off-the-shelf vision models with frozen parameters and only fine-tunes the proposed adapter. The experimental results demonstrate the superiority of our method, showing the potential for scalable IML frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。