arXiv:2409.04459cs.CRcs.CL2024-09ACL被引 9

提出线性变换水印,防止嵌入服务被改写后盗用。

WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks

  • 用线性变换对嵌入进行水印处理,提升抗改写能力。
  • 实验和理论证明可有效抵御改写攻击,水印留存率高。
  • 适合关注模型版权保护的AI服务提供商使用。

嵌入即服务(EaaS)是大型语言模型(LLM)开发者提供的服务,用于生成由LLM产生的嵌入向量。已有研究表明,EaaS容易遭受模仿攻击——攻击者通过查询得到的嵌入训练另一模型来复制底层EaaS模型。为此,研究者引入了EaaS水印以保护提供方的知识产权。本文首次表明,现有EaaS水印在攻击者对嵌入进行改写后可被移除。随后,我们提出一种新的水印技术,通过线性变换嵌入实现,并在实证和理论上证明其对改写攻击具有鲁棒性。

原文摘要 · Abstract (English)

Embeddings-as-a-Service (EaaS) is a service offered by large language model (LLM) developers to supply embeddings generated by LLMs. Previous research suggests that EaaS is prone to imitation attacks -- attacks that clone the underlying EaaS model by training another model on the queried embeddings. As a result, EaaS watermarks are introduced to protect the intellectual property of EaaS providers. In this paper, we first show that existing EaaS watermarks can be removed by paraphrasing when attackers clone the model. Subsequently, we propose a novel watermarking technique that involves linearly transforming the embeddings, and show that it is empirically and theoretically robust against paraphrasing.

嵌入服务水印版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。