arXiv:2411.09359cs.CRcs.AI2024-11EMNLP

提出语义扰动攻击,让现有版权水印失效但保持嵌入向量质量

Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark

  • 设计语义扰动攻击,利用语义不变性绕过水印验证
  • 在多个数据集上使水印识别真阳性率超95%,水印基本失效
  • 适合关注EaaS版权保护安全性的研究人员与开发者

Embedding-as-a-Service(EaaS)作为一种成功的商业模式,面临版权侵犯的严峻挑战,尤其是API滥用和模型提取攻击。已有研究提出基于后门的水印方案来保护EaaS服务版权。本文揭示了现有水印方案具有语义独立特性,并提出语义扰动攻击(SPA)。理论与实验分析表明,这种语义独立性使当前水印方案易受自适应攻击,攻击者通过语义扰动测试可绕过水印验证。在多个数据集上的大量实验显示,经SPA攻击后,水印样本的真阳性率(TPR)可达95%以上,水印失效的同时仍保持嵌入向量的高可用性。我们还讨论了潜在的防御策略。代码已公开于https://github.com/Zk4-ps/EaaS-Embedding-Watermark。

原文摘要 · Abstract (English)

Embedding-as-a-Service (EaaS) has emerged as a successful business pattern but faces significant challenges related to various forms of copyright infringement, particularly, the API misuse and model extraction attacks. Various studies have proposed backdoor-based watermarking schemes to protect the copyright of EaaS services. In this paper, we reveal that previous watermarking schemes possess semantic-independent characteristics and propose the Semantic Perturbation Attack (SPA). Our theoretical and experimental analysis demonstrate that this semantic-independent nature makes current watermarking schemes vulnerable to adaptive attacks that exploit semantic perturbations tests to bypass watermark verification. Extensive experimental results across multiple datasets demonstrate that the True Positive Rate (TPR) for identifying watermarked samples under SPA can reach up to more than 95\%, rendering watermarks ineffective while maintaining the high utility of embeddings. Furthermore, we discuss potential defense strategies to mitigate SPA. Our code is available at https://github.com/Zk4-ps/EaaS-Embedding-Watermark.

版权保护水印攻击EaaS语义扰动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。