arXiv:2606.11698cs.CRcs.AI2026-06

通过模拟盗取过程增强模型水印抗提取能力

T2S: A Rehearsal-Based Approach for Extraction-Resistant Model Watermarking

论文配图:T2S: A Rehearsal-Based Approach for Extraction-Resistant Model Watermarking
图 1 · 摘自论文原文
  • 用模拟被盗模型的损失信号指导水印优化
  • 在多种攻击下水印检测率提升显著
  • 适合保护深度学习模型知识产权

模型水印通过嵌入独特知识来保护AI模型的知识产权,产生独特的行为特征。主要技术挑战在于确保水印对各类后处理攻击的鲁棒性。模型提取攻击是最严重威胁,攻击者利用预测输出训练替代模型,非法复制原模型功能。本文提出基于回放的水印嵌入框架,以增强水印对抗模型提取攻击的鲁棒性。通过模拟提取过程,利用 extit{模拟被盗模型}在触发集上的损失作为训练信号,微调目标模型中的水印知识。该微调步骤促使水印以提高可迁移性的形式嵌入,从而增加其在被盗模型中持续存在并可被检测的概率。在多种设置下的全面实验表明,所提方法显著提升了水印对抗模型提取及后续水印移除攻击的鲁棒性。

原文摘要 · Abstract (English)

Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures. The primary technical challenge lies in ensuring watermark robustness against various post-processing attacks on the watermarked model. Model extraction attacks emerge as the most severe threat, where adversaries exploit prediction outputs to train surrogate models that illegally replicate the original model's functionality. In this work, we propose a rehearsal-based watermark embedding framework to enhance the robustness of model watermarks against model extraction attacks. By simulating the extraction process, our method leverages the loss of a \textit{simulated stolen model} on a trigger set as a training signal to fine-tune the watermark knowledge within the target model. This fine-tuning step encourages the watermark to be embedded in a way that boosts transferability, thereby increasing its chances of persisting and remaining detectable in stolen models. Comprehensive experiments conducted under diverse settings demonstrate that the proposed method significantly improves the robustness of model watermarks against both model extraction and subsequent watermark removal attacks.

模型水印抗提取知识产权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。