对比空间与隐空间水印在现代编辑工具下的脆弱性,发现二者漏洞互斥,单一方案不可靠。
The Orthogonal Vulnerabilities of Generative AI Watermarks: A Comparative Empirical Benchmark of Spatial and Latent Provenance
- 通过自动化攻击模拟,对比空间与隐空间水印的抗干扰能力
- 空间水印在图像生成编辑下逃逸率达67.47%,隐空间在裁剪下达43.20%
- 揭示当前水印标准存在系统性缺陷,需发展多域协同架构
随着开放权重生成式AI快速普及,超真实内容合成给数字信任带来严峻挑战。目前主流的隐形水印分为两类:空间域(生成后像素嵌入)与隐空间(生成前频域嵌入)。现有研究多在经典失真下评估,缺乏对现代生成式编辑工具的系统性对比。本文基于自动化攻击模拟引擎,对代表性模型RivaGAN(空间)与Tree-Ring(隐空间)在30个强度等级的几何与生成扰动下进行实证评估。提出‘对抗逃逸区’(AER)框架,衡量密码学退化与语义视觉保真度(OpenCLIP > 75.0)之间的权衡。统计分析显示(每区间n=100,MOE = ±3.92%),两类水印具有数学正交的脆弱性:空间水印在Img2Img翻译下遭遇67.47%的逃逸率;隐空间在静态裁剪下逃逸率达43.20%。证明单域水印无法应对现代攻击工具集,暴露当前数字溯源标准的系统性缺陷,为未来多域加密架构提供基础必要性依据。
原文摘要 · Abstract (English)
As open-weights generative AI rapidly proliferates, the ability to synthesize hyper-realistic media has introduced profound challenges to digital trust. Automated disinformation and AI-generated imagery have made robust digital provenance a critical cybersecurity imperative. Currently, state-of-the-art invisible watermarks operate within one of two primary mathematical manifolds: the spatial domain (post-generation pixel embedding) or the latent domain (pre-generation frequency embedding). While existing literature frequently evaluates these models against isolated, classical distortions, there is a critical lack of rigorous, comparative benchmarking against modern generative AI editing tools. In this study, we empirically evaluate two leading representative paradigms, RivaGAN (Spatial) and Tree-Ring (Latent), utilizing an automated Attack Simulation Engine across 30 intensity intervals of geometric and generative perturbations. We formalize an "Adversarial Evasion Region" (AER) framework to measure cryptographic degradation against semantic visual retention (OpenCLIP > 75.0). Our statistical analysis ($n=100$ per interval, $MOE = \pm 3.92\%$) reveals that these domains possess mutually exclusive, mathematically orthogonal vulnerabilities. Spatial watermarks experience severe cryptographic degradation under algorithmic pixel-rewriting (exhibiting a 67.47% AER evasion rate under Img2Img translation), whereas latent watermarks exhibit profound fragility against geometric misalignment (yielding a 43.20% AER evasion rate under static cropping). By proving that single-domain watermarking is fundamentally insufficient against modern adversarial toolsets, this research exposes a systemic vulnerability in current digital provenance standards and establishes the foundational exigence for future multi-domain cryptographic architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。