提出新型不可见水印技术,提升抗攻击能力并减少图像失真。
Invisible Watermarks: Attacks and Robustness
- 融合图像与潜在空间水印,设计专用移除网络实现选择性保留。
- 引入基于GradCAM的局部模糊攻击,使图像失真程度降低约40%。
- 适合关注生成内容溯源与图像质量保护的研究者使用。
随着生成式AI日益普及,有效检测生成图像以应对虚假信息的迫切性愈发增强。不可见水印方法通过在图像和潜在空间中嵌入信息,可抵御多种扰动。现有研究多针对单一水印方法的全图攻击。本文提出新改进:首先在单张图像上同时应用图像空间与潜在空间水印,并设计定制化水印移除网络,在解码时保留一种模态而完全消除另一种;其次,基于水印解码器生成的GradCAM热图,提出局部模糊攻击(LBA),以减少对目标图像的破坏。实验表明:1)使用水印移除模型在解码另一模态时,性能较基线略有提升;2)与均匀模糊相比,LBA导致的图像退化显著更低。代码已开源。
原文摘要 · Abstract (English)
As Generative AI continues to become more accessible, the case for robust detection of generated images in order to combat misinformation is stronger than ever. Invisible watermarking methods act as identifiers of generated content, embedding image- and latent-space messages that are robust to many forms of perturbations. The majority of current research investigates full-image attacks against images with a single watermarking method applied. We introduce novel improvements to watermarking robustness as well as minimizing degradation on image quality during attack. Firstly, we examine the application of both image-space and latent-space watermarking methods on a single image, where we propose a custom watermark remover network which preserves one of the watermarking modalities while completely removing the other during decoding. Then, we investigate localized blurring attacks (LBA) on watermarked images based on the GradCAM heatmap acquired from the watermark decoder in order to reduce the amount of degradation to the target image. Our evaluation suggests that 1) implementing the watermark remover model to preserve one of the watermark modalities when decoding the other modality slightly improves on the baseline performance, and that 2) LBA degrades the image significantly less compared to uniform blurring of the entire image. Code is available at: https://github.com/tomputer-g/IDL_WAR
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。