评估大模型隐空间水印在复制和移除攻击下的安全性
Evaluation of Security of ML-based Watermarking: Copy and Removal Attacks
- 用对抗嵌入技术在大模型隐空间中加水印
- 发现水印易被复制和移除,安全防御不足
- 适合关注AI内容版权保护的研究者
从真实世界或AI生成的数字内容激增,亟需版权保护、溯源与数据可信验证方法。数字水印是解决此类问题的关键技术,其发展历经手工设计、自编码器和基础模型三阶段。尽管系统鲁棒性已有充分研究,但针对对抗攻击的安全性仍待深入。本文评估基于基础模型隐空间的数字水印系统在对抗嵌入技术下的安全性,通过一系列实验分析其在复制和移除攻击下的表现,揭示潜在漏洞。所有实验代码与结果公开于 https://github.com/vkinakh/ssl-watermarking-attacks。
原文摘要 · Abstract (English)
The vast amounts of digital content captured from the real world or AI-generated media necessitate methods for copyright protection, traceability, or data provenance verification. Digital watermarking serves as a crucial approach to address these challenges. Its evolution spans three generations: handcrafted, autoencoder-based, and foundation model based methods. While the robustness of these systems is well-documented, the security against adversarial attacks remains underexplored. This paper evaluates the security of foundation models' latent space digital watermarking systems that utilize adversarial embedding techniques. A series of experiments investigate the security dimensions under copy and removal attacks, providing empirical insights into these systems' vulnerabilities. All experimental codes and results are available at https://github.com/vkinakh/ssl-watermarking-attacks .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。