提出首个联合音视频生成的水印框架,防止音频替换攻击。
mAVE: A Watermark for Joint Audio-Visual Generation Models
- 在初始化时加密绑定音视频潜在表示,无需微调。
- 对换脸攻击实现超99%绑定完整性,性能无损失。
- 适合需版权保护的音视频生成系统使用。
随着联合音视频生成模型广泛应用于商业场景,嵌入水印成为保护厂商版权和确保内容溯源的关键。然而,现有技术因将模态视为独立实体,存在严重绑定漏洞。攻击者通过替换真实音频为恶意深度伪造音频,同时保留水印视频,导致现有检测器(依赖独立验证 $Video_{wm}\vee Audio_{wm}$)错误认证篡改内容,错误归因于原厂商并严重损害其声誉。为此,我们提出 mAVE(Manifold Audio-Visual Entanglement),首个专为联合架构设计的水印框架。mAVE 在初始化时加密绑定音视频潜在表示,无需微调,通过逆变换采样定义合法纠缠流形。在 LTX-2、MOVA 等先进模型上的实验表明,mAVE 实现性能无损,并对换脸攻击提供指数级安全边界。绑定完整性超过99%,为厂商版权提供强加密防护。
原文摘要 · Abstract (English)
As Joint Audio-Visual Generation Models see widespread commercial deployment, embedding watermarks has become essential for protecting vendor copyright and ensuring content provenance. However, existing techniques suffer from an architectural mismatch by treating modalities as decoupled entities, exposing a critical Binding Vulnerability. Adversaries exploit this via Swap Attacks by replacing authentic audio with malicious deepfakes while retaining the watermarked video. Because current detectors rely on independent verification ($Video_{wm}\vee Audio_{wm}$), they incorrectly authenticate the manipulated content, falsely attributing harmful media to the original vendor and severely damaging their reputation. To address this, we propose mAVE (Manifold Audio-Visual Entanglement), the first watermarking framework natively designed for joint architectures. mAVE cryptographically binds audio and video latents at initialization without fine-tuning, defining a Legitimate Entanglement Manifold via Inverse Transform Sampling. Experiments on state-of-the-art models (LTX-2, MOVA) demonstrate that mAVE guarantees performance-losslessness and provides an exponential security bound against Swap Attacks. Achieving near-perfect binding integrity ($>99\%$), mAVE offers a robust cryptographic defense for vendor copyright.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。