arXiv:2603.00194cs.CVcs.AI2026-03

为文生视频模型设计高鲁棒水印,防篡改且不失真。

SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models

  • 用打乱密钥生成帧级密钥,抗帧序错乱
  • 动态调整注意力,抵抗帧间压缩等时序失真
  • 适合内容版权保护与真实性的可信验证

文本到视频生成模型的兴起引发了对内容真实性、版权保护和恶意滥用的担忧。水印是监管此类生成内容的有效手段,其中高保真度和强鲁棒性尤为关键。现有生成式图像水印方法通过水印信息和伪随机密钥控制初始采样噪声,实现无损嵌入,但直接应用于视频会面临两个核心问题:现有设计依赖帧与帧相关伪随机二进制序列的严格对齐,一旦对齐被破坏,水印提取将不可靠;视频特有的帧间压缩等失真会显著降低水印可靠性。为此,我们提出 SKeDA,一种专为文生视频扩散模型设计的生成式水印框架。SKeDA 包含两个组件:(1) 基于打乱密钥的分布保持采样(SKe),使用单一基础伪随机二进制序列进行水印加密,并通过置换生成帧级加密序列,将水印提取从敏感同步的序列解码转变为容忍置换的集合级聚合,大幅提高对帧重排序和丢失的鲁棒性;(2) 差分注意力(DA),计算帧间差异并动态调整提取过程中的注意力权重,增强对时序失真的鲁棒性。大量实验表明,SKeDA 在保持高视频生成质量的同时,显著提升了水印鲁棒性。

原文摘要 · Abstract (English)

The rise of text-to-video generation models has raised growing concerns over content authenticity, copyright protection, and malicious misuse. Watermarking serves as an effective mechanism for regulating such AI-generated content, where high fidelity and strong robustness are particularly critical. Recent generative image watermarking methods provide a promising foundation by leveraging watermark information and pseudo-random keys to control the initial sampling noise, enabling lossless embedding. However, directly extending these techniques to videos introduces two key limitations: Existing designs implicitly rely on strict alignment between video frames and frame-dependent pseudo-random binary sequences used for watermark encryption. Once this alignment is disrupted, subsequent watermark extraction becomes unreliable; and Video-specific distortions, such as inter-frame compression, significantly degrade watermark reliability. To address these issues, we propose SKeDA, a generative watermarking framework tailored for text-to-video diffusion models. SKeDA consists of two components: (1) Shuffle-Key-based Distribution-preserving Sampling (SKe) employs a single base pseudo-random binary sequence for watermark encryption and derives frame-level encryption sequences through permutation. This design transforms watermark extraction from synchronization-sensitive sequence decoding into permutation-tolerant set-level aggregation, substantially improving robustness against frame reordering and loss; and (2) Differential Attention (DA), which computes inter-frame differences and dynamically adjusts attention weights during extraction, enhancing robustness against temporal distortions. Extensive experiments demonstrate that SKeDA preserves high video generation quality and watermark robustness.

视频生成水印技术扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。