arXiv:2501.09328cs.CRcs.AI2025-01

无需重训练的水印框架,有效抵御模型窃取攻击

Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks

  • 不依赖训练,通过多步传播机制嵌入水印
  • 验证时仅需现有方法2%的查询次数
  • 适合部署后快速添加水印的模型所有者

可触发水印技术使模型所有者能够对抗模型提取攻击。然而,现有大多数方法需要额外训练,限制了部署后的灵活性,且缺乏清晰的理论基础,易受自适应攻击。本文提出 Neural Honeytrace,一种无需重训练的即插即用式水印框架。我们从信息论角度重新定义水印传输机制,设计了一种免训练的多步传输策略,利用后门学习的长尾效应实现高效且鲁棒的水印嵌入。大量实验表明,Neural Honeytrace 将最坏情况下的基于 t 检验的产权验证所需平均查询次数降低至现有方法的 2%,同时训练成本为零。

原文摘要 · Abstract (English)

Triggerable watermarking enables model owners to assert ownership against model extraction attacks. However, most existing approaches require additional training, which limits post-deployment flexibility, and the lack of clear theoretical foundations makes them vulnerable to adaptive attacks. In this paper, we propose Neural Honeytrace, a plug-and-play watermarking framework that operates without retraining. We redefine the watermark transmission mechanism from an information perspective, designing a training-free multi-step transmission strategy that leverages the long-tailed effect of backdoor learning to achieve efficient and robust watermark embedding. Extensive experiments demonstrate that Neural Honeytrace reduces the average number of queries required for a worst-case t-test-based ownership verification to as low as $2\%$ of existing methods, while incurring zero training cost.

水印技术模型安全后门防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。