用对抗强化学习在潜在空间嵌入水印,提升伪造视频检测的鲁棒性与敏感性。
DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Adversarial Reinforcement Learning

- 在图像潜在空间中学习嵌入水印,捕捉高层语义信息
- 在挑战性篡改下,准确率比现有方法高4.5%以上
- 适合需要高可靠伪造检测的媒体验证场景
生成式AI的快速发展使深度伪造技术愈发逼真,对执法和公众信任构成严峻挑战。现有被动检测方法依赖特定伪造痕迹,难以泛化至新型伪造。主动水印检测可应对这一问题,但常难以兼顾对正常失真(如压缩、旋转)的鲁棒性与对恶意篡改的敏感性。本文提出一种基于高维潜在空间表示与对抗强化学习(ARL)的新框架。设计一个可学习的水印嵌入器,在潜在空间中编码信息,同时实现对消息嵌入与提取的精确控制。通过对抗攻击代理动态生成良性与恶意篡改数据,引导水印模块在鲁棒性与脆弱性之间取得最优平衡。在CelebA与CelebA-HQ基准上的全面评估显示,本方法在挑战性篡改场景下持续优于现有先进方法:在CelebA上提升超4.5%,在CelebA-HQ上提升超5.3%。
原文摘要 · Abstract (English)
Rapid advances in generative AI have led to increasingly realistic deepfakes, posing growing challenges for law enforcement and public trust. Existing passive deepfake detectors struggle to keep pace, largely due to their dependence on specific forgery artifacts, which limits their ability to generalize to new deepfake types. Proactive deepfake detection using watermarks has emerged to address the challenge of identifying high-quality synthetic media. However, these methods often struggle to balance robustness against benign distortions with sensitivity to malicious tampering. This paper introduces a novel deep learning framework that harnesses high-dimensional latent space representations and the Adversarial Reinforcement Learning (ARL) paradigm to develop a robust and adaptive watermarking approach. Specifically, we develop a learnable watermark embedder that operates in the latent space, capturing high-level image semantics, while offering precise control over message encoding and extraction. The ARL paradigm empowers the learnable watermarking module to pursue an optimal balance between robustness and fragility. This is achieved through interaction with a dynamic curriculum of benign and malicious image manipulations simulated by an adversarial attacker agent. Comprehensive evaluations on the CelebA and CelebA-HQ benchmarks reveal that our method consistently outperforms state-of-the-art approaches, achieving improvements of over 4.5% on CelebA and more than 5.3% on CelebA-HQ under challenging manipulation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。