arXiv:2505.12667cs.CV2025-05NeurIPS被引 9

首个在视频生成中嵌入图形水印的框架,提升版权保护可靠性

Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking

  • 通过分块匹配将水印与视频帧视觉相似区域对齐
  • 3D小波增强的Mamba模型实现水印跨帧融合,保真度超90%
  • 首次将状态空间模型用于水印,适合内容创作者与平台方

生成式视频模型的爆炸式增长加剧了对AI生成内容版权保护的需求。尽管隐形生成水印在图像合成中已广泛应用,但在视频生成领域仍研究不足。为此,我们提出Safe-Sora,首个直接在视频生成过程中嵌入图形水印的框架。基于水印与载体内容视觉相似性影响嵌入效果的观察,我们设计了从粗到细的自适应匹配机制:将水印图像划分为块,分配给最相似的视频帧,并进一步定位到最优空间区域以实现无缝嵌入。为支持跨视频帧的时空水印融合,我们开发了一种基于3D小波变换增强的Mamba架构,结合新颖的时空局部扫描策略,有效建模长距离依赖关系。据我们所知,这是首次将状态空间模型应用于水印任务,开辟了高效鲁棒水印保护的新路径。大量实验表明,Safe-Sora在视频质量、水印保真度和鲁棒性方面均达到当前最优水平,主要归功于上述设计。代码已公开于https://github.com/Sugewud/Safe-Sora。

原文摘要 · Abstract (English)

The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, invisible generative watermarking remains largely underexplored in video generation. To address this gap, we propose Safe-Sora, the first framework to embed graphical watermarks directly into the video generation process. Motivated by the observation that watermarking performance is closely tied to the visual similarity between the watermark and cover content, we introduce a hierarchical coarse-to-fine adaptive matching mechanism. Specifically, the watermark image is divided into patches, each assigned to the most visually similar video frame, and further localized to the optimal spatial region for seamless embedding. To enable spatiotemporal fusion of watermark patches across video frames, we develop a 3D wavelet transform-enhanced Mamba architecture with a novel spatiotemporal local scanning strategy, effectively modeling long-range dependencies during watermark embedding and retrieval. To the best of our knowledge, this is the first attempt to apply state space models to watermarking, opening new avenues for efficient and robust watermark protection. Extensive experiments demonstrate that Safe-Sora achieves state-of-the-art performance in terms of video quality, watermark fidelity, and robustness, which is largely attributed to our proposals. Code is publicly available at https://github.com/Sugewud/Safe-Sora

视频生成水印技术生成模型版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。