通过感知内在重要性,自适应优化伪造图像检测特征表示。
Adaptive Forensic Feature Refinement via Intrinsic Importance Perception

- 基于视觉基础模型,自适应识别关键特征层以提升判别力。
- 在保持预训练结构稳定的同时,实现任务特异性参数更新。
- 适合需要跨源泛化能力的伪造图像检测场景。
随着生成模型和多模态编辑技术的快速发展,合成图像检测(SID)面临的关键挑战在于对未知生成源的跨分布泛化能力。近年来,通过大规模图文对齐预训练获得丰富视觉先验的视觉基础模型(VFM)成为提升SID泛化能力的有前景路径。然而,现有基于VFM的方法适应策略仍较粗粒度,通常直接使用最终层表征或简单融合多层特征,缺乏对可迁移伪造线索最优表征层次的显式建模。尽管直接微调VFM可增强任务适应性,但可能破坏支持开集泛化的跨模态预训练结构。为解决这一任务特异性矛盾,本文将VFM适配重新构想为联合优化问题:既要识别更适于承载伪造判别信息的代表性层级,又要约束任务知识注入对预训练结构的扰动。基于此,提出以内在重要性感知为核心的I2P框架。I2P首先自适应识别对SID最具判别性的关键层表征,随后在低敏感度参数子空间内约束任务驱动的参数更新,从而在提升任务特异性的同时最大程度保留预训练表征的可迁移结构。
原文摘要 · Abstract (English)
With the rapid development of generative models and multimodal content editing technologies, the key challenge faced by synthetic image detection (SID) lies in cross-distribution generalization to unknown generation sources. In recent years, visual foundation models (VFM), which acquire rich visual priors through large scale image-text alignment pretraining, have become a promising technical route for improving the generalization ability of SID. However, existing VFM-based methods remain relatively coarse-grained in their adaptation strategies. They typically either directly use the final layer representations of VFM or simply fuse multi layer features, lacking explicit modeling of the optimal representational hierarchy for transferable forgery cues. Meanwhile, although directly fine-tuning VFM can enhance task adaptation, it may also damage the cross-modal pretrained structure that supports open-set generalization. To address this task specific tension, we reformulate VFM adaptation for SID as a joint optimization problem: it is necessary both to identify the critical representational layer that is more suitable for carrying forgery discriminative information and to constrain the disturbance caused by task knowledge injection to the pretrained structure. Based on this, we propose I2P, an SID framework centered on intrinsic importance perception. I2P first adaptively identifies the critical layer representations that are most discriminative for SID, and then constrains task-driven parameter updates within a low sensitivity parameter subspace, thereby improving task specificity while preserving the transferable structure of pretrained representations as much as possible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。