提出抗伪造攻击的智能水印技术,提升模型版权保护可靠性
Agentic Copyright Watermarking against Adversarial Evidence Forgery with Purification-Agnostic Curriculum Proxy Learning
- 用哈希构建自验证黑盒水印协议,无需模型内部信息
- 发现对抗扰动可伪造水印证据,提出净化无关的渐进学习防御
- 适合模型版权保护需求方,尤其关注安全性的开发者
随着AI代理在各领域的普及,模型所有权保护变得至关重要,因其开发投入巨大。未经授权使用和非法分发严重威胁知识产权,亟需有效的版权保护措施。模型水印已成为关键手段,通过在模型中嵌入所有权信息以在版权争议中主张权利。本文提出多项贡献:基于哈希技术的自验证黑盒水印协议;研究利用对抗扰动进行证据伪造攻击;提出包含净化步骤的防御机制以应对攻击;以及一种不依赖净化过程的渐进式代理学习方法,以增强水印鲁棒性与模型性能。实验结果表明,这些方法有效提升了水印模型的安全性、可靠性和性能。
原文摘要 · Abstract (English)
With the proliferation of AI agents in various domains, protecting the ownership of AI models has become crucial due to the significant investment in their development. Unauthorized use and illegal distribution of these models pose serious threats to intellectual property, necessitating effective copyright protection measures. Model watermarking has emerged as a key technique to address this issue, embedding ownership information within models to assert rightful ownership during copyright disputes. This paper presents several contributions to model watermarking: a self-authenticating black-box watermarking protocol using hash techniques, a study on evidence forgery attacks using adversarial perturbations, a proposed defense involving a purification step to counter adversarial attacks, and a purification-agnostic curriculum proxy learning method to enhance watermark robustness and model performance. Experimental results demonstrate the effectiveness of these approaches in improving the security, reliability, and performance of watermarked models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。