arXiv:2605.16720cs.CVcs.LG2026-05

用可学习的组合攻击训练水印,显著提升抗攻击能力。

Compositional Adversarial Training for Robust Visual Watermarking

论文配图:Compositional Adversarial Training for Robust Visual Watermarking
图 1 · 摘自论文原文
  • 设计可微分的序列攻击选择机制,动态生成组合攻击
  • 在单步和组合攻击下,水印容量最高提升63.5%
  • 适合需要强鲁棒性的视频/图像水印应用

鲁棒水印通常通过随机后处理增强进行训练,但随机采样难以覆盖真实攻击链的组合空间,且极少遇到实际导致检测失败的罕见组合,导致训练不稳定且样本效率低。本文将水印鲁棒性建模为组合变换结构空间上的极小-极大问题,提出可组合对抗训练(CAT)框架。CAT是一种即插即用的方案,能学习一个可微分的序列对抗者,在每一步观察当前带水印图像并选择攻击族以最大化消息恢复破坏。该方法结合直通Gumbel-Softmax攻击选择与熵正则化,使反向传播端到端可微,实现跨攻击族梯度信息聚合,从而加速收敛并避免陷入单一攻击模式。我们在Post-generation水印VideoSeal 0.0、VideoSeal 1.0、PixelSeal及In-generation WMAR上评估,涵盖单步和双步攻击设置,以及分布内与多个分布外图像/视频基准。结果表明,CAT始终优于相同增强预算下的随机增强基线,尤其在复杂组合攻击和OoD评估中表现突出:单步攻击下水印容量最高提升63.5%,组合攻击下提升13.0%;自回归设置中,对困难几何变换的TPR@FPR=1%平均提升12%。结果证明,针对自适应组合对抗者训练比独立随机扰动更有利于提升视觉水印鲁棒性。

原文摘要 · Abstract (English)

Robust watermarking is typically trained with random post-processing augmentation, but random sampling under-covers the combinatorial space of realistic attack pipelines and rarely encounters the rare compositions that actually break detection. This leads to unstable training and poor sample efficiency. We instead formulate watermark robustness as a min-max problem over a structured space of compositional transformations. We propose Compositional Adversarial Training (CAT), a plug-in framework that learns a sequential differentiable adversary that observes the current watermarked image and selects an attack family at each step to maximally disrupt message recovery. CAT combines a straight-through Gumbel-Softmax attack selection with entropy regularization, allowing the backward pass to be end-to-end differentiable and aggregate gradient information across attack families, yielding faster, smoother convergence without collapsing to a single attack mode. We evaluate CAT on post-generation watermarks VideoSeal 0.0, VideoSeal 1.0, and PixelSeal and in-generation WMAR under both single-step and two-step attack suites, on in-distribution and multiple out-of-distribution image and video benchmarks. CAT consistently outperforms random-augmentation baselines trained with the same augmentation budget, with the largest gains on hard composed attacks and OOD evaluations; improving overall watermark capacity by up to $63.5\%$ in the single-step attack setting and $13.0\%$ in the compositional setting. In the autoregressive setting, CAT improves the TPR@FPR$=1\%$ by $12\%$ on average on difficult geometric transformations. These results show that robust visual watermarking benefits from training against adaptive compositional adversaries rather than independent random corruptions.

水印对抗训练鲁棒性组合攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。