通过噪声与边缘监督,精准定位深伪和浅伪图像的篡改区域。
A Noise and Edge extraction-based dual-branch method for Shallowfake and Deepfake Localization
- 双分支结构分别提取噪声特征和RGB语义特征
- 在多个数据集上达到99% AUC,显著优于现有方法
- 适合需要高精度篡改定位的多媒体安全研究者
多媒体可信度评估日益依赖先进的图像篡改定位(IML)技术,推动了该领域的发展。有效的篡改模型需提取篡改区与真实区之间的非语义差异特征以利用伪造痕迹,这要求直接比较两区域。现有模型多采用手工特征、卷积神经网络(CNN)或二者结合的方法。手工特征需预设篡改类型,限制其泛化能力;而CNN捕捉语义信息,对伪造痕迹敏感性不足。为此,本文提出一种双分支模型,融合人工设计的噪声特征与传统CNN特征。该模型采用双分支策略:一支整合噪声特性,另一支利用层级化的ConvNext模块提取RGB特征。此外,引入边缘监督损失,精准获取篡改边界信息。同时,通过特征增强模块优化特征表示。在浅伪数据集(CASIA、COVERAGE、COLUMBIA、NIST16)与深伪数据集Faceforensics++(FF++)上进行充分测试,验证其卓越的特征提取能力与性能优势,AUC高达99%,显著超越现有最先进模型。
原文摘要 · Abstract (English)
The trustworthiness of multimedia is being increasingly evaluated by advanced Image Manipulation Localization (IML) techniques, resulting in the emergence of the IML field. An effective manipulation model necessitates the extraction of non-semantic differential features between manipulated and legitimate sections to utilize artifacts. This requires direct comparisons between the two regions.. Current models employ either feature approaches based on handcrafted features, convolutional neural networks (CNNs), or a hybrid approach that combines both. Handcrafted feature approaches presuppose tampering in advance, hence restricting their effectiveness in handling various tampering procedures, but CNNs capture semantic information, which is insufficient for addressing manipulation artifacts. In order to address these constraints, we have developed a dual-branch model that integrates manually designed feature noise with conventional CNN features. This model employs a dual-branch strategy, where one branch integrates noise characteristics and the other branch integrates RGB features using the hierarchical ConvNext Module. In addition, the model utilizes edge supervision loss to acquire boundary manipulation information, resulting in accurate localization at the edges. Furthermore, this architecture utilizes a feature augmentation module to optimize and refine the presentation of attributes. The shallowfakes dataset (CASIA, COVERAGE, COLUMBIA, NIST16) and deepfake dataset Faceforensics++ (FF++) underwent thorough testing to demonstrate their outstanding ability to extract features and their superior performance compared to other baseline models. The AUC score achieved an astounding 99%. The model is superior in comparison and easily outperforms the existing state-of-the-art (SoTA) models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。