arXiv:2410.03537cs.LGcs.AI2024-10ICLR被引 17

用大模型水印检测外部数据在RAG中被滥用,提供可证明的保护方案。

Ward: Provable RAG Dataset Inference via LLM Watermarks

  • 基于大模型水印设计新方法,实现对RAG数据滥用的精准检测。
  • 在真实场景下测试,准确率高于基线方法,且查询效率更高。
  • 适合关注数据版权保护的研究者和开发者使用。

RAG使大模型能便捷地整合外部数据,引发数据所有者对其内容被未经授权使用的担忧。当前对这种滥用行为的检测研究仍不充分,现有数据集和方法难以适配该问题。我们采取三步举措填补这一空白:首先将该问题形式化为(黑盒)RAG数据集推断(RAG-DI);随后构建一个面向真实评估的新型数据集及一组基线方法;最后提出Ward,一种基于大模型水印的RAG-DI方法,为数据所有者提供关于其数据在RAG语料中被滥用的严格统计保障。Ward在各项指标上均优于所有基线,展现出更高的准确性、更优的查询效率和更强的鲁棒性。本工作为未来RAG-DI研究奠定基础,并凸显大模型水印在此问题上的巨大潜力。

原文摘要 · Abstract (English)

RAG enables LLMs to easily incorporate external data, raising concerns for data owners regarding unauthorized usage of their content. The challenge of detecting such unauthorized usage remains underexplored, with datasets and methods from adjacent fields being ill-suited for its study. We take several steps to bridge this gap. First, we formalize this problem as (black-box) RAG Dataset Inference (RAG-DI). We then introduce a novel dataset designed for realistic benchmarking of RAG-DI methods, alongside a set of baselines. Finally, we propose Ward, a method for RAG-DI based on LLM watermarks that equips data owners with rigorous statistical guarantees regarding their dataset's misuse in RAG corpora. Ward consistently outperforms all baselines, achieving higher accuracy, superior query efficiency and robustness. Our work provides a foundation for future studies of RAG-DI and highlights LLM watermarks as a promising approach to this problem.

RAG水印数据安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。