通过强耦合水印机制,让盗版模型无法移除版权标记。
DeepTracer: Tracing Stolen Model via Deep Coupled Watermarks
- 设计水印样本生成与同类别耦合损失,使水印与主任务深度绑定。
- 在多个数据集上对齐攻击下仍保持90%以上水印检测准确率。
- 适合模型版权保护场景,尤其对抗模型窃取攻击。
模型水印技术通过构建特定输入输出对将水印信息嵌入被保护模型以声明所有权。然而,现有水印方法在面对模型窃取攻击时极易被移除,导致模型所有者难以有效验证被盗模型的版权。本文分析了当前水印方法在模型窃取场景下失效的根本原因,并探索可行解决方案。具体而言,提出一种鲁棒水印框架DeepTracer,采用新颖的水印样本构造方法和同类别耦合损失约束,使水印任务与主任务形成高耦合关系,从而迫使攻击者在窃取主任务功能时不可避免地学习隐藏的水印任务。此外,设计了一种有效的水印样本过滤机制,精炼用于模型所有权验证的关键水印样本,提升水印可靠性。大量实验在多个数据集和模型上表明,该方法在防御多种模型窃取攻击及水印攻击方面均优于现有方法,达到新的最先进效果与鲁棒性。
原文摘要 · Abstract (English)
Model watermarking techniques can embed watermark information into the protected model for ownership declaration by constructing specific input-output pairs. However, existing watermarks are easily removed when facing model stealing attacks, and make it difficult for model owners to effectively verify the copyright of stolen models. In this paper, we analyze the root cause of the failure of current watermarking methods under model stealing scenarios and then explore potential solutions. Specifically, we introduce a robust watermarking framework, DeepTracer, which leverages a novel watermark samples construction method and a same-class coupling loss constraint. DeepTracer can incur a high-coupling model between watermark task and primary task that makes adversaries inevitably learn the hidden watermark task when stealing the primary task functionality. Furthermore, we propose an effective watermark samples filtering mechanism that elaborately select watermark key samples used in model ownership verification to enhance the reliability of watermarks. Extensive experiments across multiple datasets and models demonstrate that our method surpasses existing approaches in defending against various model stealing attacks, as well as watermark attacks, and achieves new state-of-the-art effectiveness and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。