用Swin Transformer提升图像水印抗几何攻击能力
RoWSFormer: A Robust Watermarking Framework with Swin Transformer for Enhanced Geometric Attack Resilience
- 基于Swin Transformer设计新型水印编码解码模块,捕捉全局信息
- 对几何攻击(旋转/缩放)的PSNR提升超6 dB,提取准确率超97%
- 适合需要高鲁棒性水印的应用,如版权保护与内容认证
近年来,基于深度学习的数字水印技术受到广泛关注。为实现图像水印的不可感知性和鲁棒性,现有方法多采用卷积神经网络构建水印框架。然而,由于卷积网络难以捕捉全局和长距离关系,其在几何攻击下表现不佳。为此,本文提出基于Swin Transformer的鲁棒水印框架RoWSFormer。具体地,设计了局部-通道增强型Swin Transformer模块作为编码器和解码器的核心,利用自注意力机制捕获全局长程信息,显著提升对几何失真的适应能力。此外,构建频率增强型Transformer模块以提取频域特征,进一步增强鲁棒性。实验表明,相较于现有最优方法,RoWSFormer在多数非几何攻击下,PSNR提升3 dB且保持相同提取准确率;在几何攻击(如旋转、缩放、仿射变换)下,PSNR提升超过6 dB,提取准确率超过97%。
原文摘要 · Abstract (English)
In recent years, digital watermarking techniques based on deep learning have been widely studied. To achieve both imperceptibility and robustness of image watermarks, most current methods employ convolutional neural networks to build robust watermarking frameworks. However, despite the success of CNN-based watermarking models, they struggle to achieve robustness against geometric attacks due to the limitations of convolutional neural networks in capturing global and long-range relationships. To address this limitation, we propose a robust watermarking framework based on the Swin Transformer, named RoWSFormer. Specifically, we design the Locally-Channel Enhanced Swin Transformer Block as the core of both the encoder and decoder. This block utilizes the self-attention mechanism to capture global and long-range information, thereby significantly improving adaptation to geometric distortions. Additionally, we construct the Frequency-Enhanced Transformer Block to extract frequency domain information, which further strengthens the robustness of the watermarking framework. Experimental results demonstrate that our RoWSFormer surpasses existing state-of-the-art watermarking methods. For most non-geometric attacks, RoWSFormer improves the PSNR by 3 dB while maintaining the same extraction accuracy. In the case of geometric attacks (such as rotation, scaling, and affine transformations), RoWSFormer achieves over a 6 dB improvement in PSNR, with extraction accuracy exceeding 97\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。