用不可见水印防深度伪造,能认出篡改却不怕常规处理。
Social Media Authentication and Combating Deepfakes using Semi-fragile Invisible Image Watermarking
- 设计对抗网络架构,让水印对人脸篡改敏感但抗正常图像操作。
- 64位秘密信息嵌入后,正常处理下恢复准确率高,深度伪造后无法恢复。
- 对白盒和黑盒水印移除攻击均有强抵抗力,适合媒体真伪验证场景。
随着深度生成模型在图像与视频合成中的快速发展,深度伪造和篡改媒体引发了严重社会问题。传统机器学习检测方法难以应对不断演进的生成技术,且易受对抗攻击。相比之下,不可见图像水印作为主动防御手段,可通过验证嵌入图像像素中的隐秘信息实现媒体认证。现有少数不可见水印技术易受基本图像处理和水印移除攻击影响。为此,本文提出一种半脆弱图像水印技术,将不可见的秘密消息嵌入真实图像中用于媒体认证。所提框架对人脸篡改或篡改操作敏感,但对良性图像处理操作及水印移除攻击具有鲁棒性。该特性通过独特的架构实现:包含判别器与对抗网络分别保障图像质量与抗移除能力,结合主干编码器-解码器与判别网络。在主流人脸深度伪造数据集上的充分实验表明,本模型可嵌入64位秘密信息为不可感知水印,在施加良性图像处理时可高精度恢复,而在面对未见过的深度伪造操作时则无法恢复。此外,该水印技术对多种白盒与黑盒水印移除攻击表现出高度鲁棒性,达到当前最优性能。
原文摘要 · Abstract (English)
With the significant advances in deep generative models for image and video synthesis, Deepfakes and manipulated media have raised severe societal concerns. Conventional machine learning classifiers for deepfake detection often fail to cope with evolving deepfake generation technology and are susceptible to adversarial attacks. Alternatively, invisible image watermarking is being researched as a proactive defense technique that allows media authentication by verifying an invisible secret message embedded in the image pixels. A handful of invisible image watermarking techniques introduced for media authentication have proven vulnerable to basic image processing operations and watermark removal attacks. In response, we have proposed a semi-fragile image watermarking technique that embeds an invisible secret message into real images for media authentication. Our proposed watermarking framework is designed to be fragile to facial manipulations or tampering while being robust to benign image-processing operations and watermark removal attacks. This is facilitated through a unique architecture of our proposed technique consisting of critic and adversarial networks that enforce high image quality and resiliency to watermark removal efforts, respectively, along with the backbone encoder-decoder and the discriminator networks. Thorough experimental investigations on SOTA facial Deepfake datasets demonstrate that our proposed model can embed a $64$-bit secret as an imperceptible image watermark that can be recovered with a high-bit recovery accuracy when benign image processing operations are applied while being non-recoverable when unseen Deepfake manipulations are applied. In addition, our proposed watermarking technique demonstrates high resilience to several white-box and black-box watermark removal attacks. Thus, obtaining state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。