arXiv:2501.05249cs.CRcs.AI2025-01被引 10

为大模型检索增强生成系统设计黑盒水印,防抄袭且抗干扰。

RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language Models

  • 通过多模型协作生成知识型水印并注入RAG系统。
  • 在多种任务中检测盗用率超95%,抗改写、删减等攻击。
  • 适合保护医疗、法律等敏感领域RAG系统的知识产权。

近年来,检索增强生成(RAG)在领域特定、知识密集和隐私敏感任务中广泛应用,显著提升大语言模型(LLM)性能。然而,攻击者可能窃取有价值的RAG系统并部署或商业化,亟需知识产权(IP)侵权检测机制。现有水印方法多针对文本或关系数据库,依赖白盒访问,无法适用于RAG的知识库;且部署后的LLM后处理常破坏文本水印信息。为此,本文提出一种新型黑盒“知识水印”方法RAG-WM,通过水印生成器、影子LLM&RAG与水印判别器组成的多模型交互框架,基于水印实体-关系三元组生成水印内容并注入目标RAG。在四个基准大模型上,对三个领域特定及两个隐私敏感任务进行评估,结果表明RAG-WM可有效检测各类部署的盗用RAG系统。此外,该方法对重述、无关内容删除、知识插入和知识扩展攻击具有鲁棒性,且能规避主流水印检测手段,展现出在RAG系统知识产权保护中的广阔应用前景。

原文摘要 · Abstract (English)

In recent years, tremendous success has been witnessed in Retrieval-Augmented Generation (RAG), widely used to enhance Large Language Models (LLMs) in domain-specific, knowledge-intensive, and privacy-sensitive tasks. However, attackers may steal those valuable RAGs and deploy or commercialize them, making it essential to detect Intellectual Property (IP) infringement. Most existing ownership protection solutions, such as watermarks, are designed for relational databases and texts. They cannot be directly applied to RAGs because relational database watermarks require white-box access to detect IP infringement, which is unrealistic for the knowledge base in RAGs. Meanwhile, post-processing by the adversary's deployed LLMs typically destructs text watermark information. To address those problems, we propose a novel black-box "knowledge watermark" approach, named RAG-WM, to detect IP infringement of RAGs. RAG-WM uses a multi-LLM interaction framework, comprising a Watermark Generator, Shadow LLM & RAG, and Watermark Discriminator, to create watermark texts based on watermark entity-relationship tuples and inject them into the target RAG. We evaluate RAG-WM across three domain-specific and two privacy-sensitive tasks on four benchmark LLMs. Experimental results show that RAG-WM effectively detects the stolen RAGs in various deployed LLMs. Furthermore, RAG-WM is robust against paraphrasing, unrelated content removal, knowledge insertion, and knowledge expansion attacks. Lastly, RAG-WM can also evade watermark detection approaches, highlighting its promising application in detecting IP infringement of RAG systems.

RAG水印知识产权黑盒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。