用统计检验方法检测大模型是否盗用他人生成数据
A Statistical Hypothesis Testing Framework for Data Misappropriation Detection in Large Language Models
- 给版权数据加水印,转为假设检验问题
- 可控制误报率和漏报率,理论最优性已证明
- 适合关注模型版权与数据安全的研究者
大型语言模型(LLM)近年来迅速流行,但其训练过程引发了重大的隐私与法律争议,尤其涉及在未授权或未署名的情况下,将受版权保护的内容纳入训练数据中,这属于数据盗用的范畴。本文聚焦于一种特定的数据盗用检测问题:判断某个大模型是否包含另一个大模型生成的数据。我们通过在受版权保护的训练数据中嵌入水印,将数据盗用检测问题建模为假设检验。提出一个通用的统计检验框架,构建检验统计量,确定最优拒绝阈值,并显式控制第一类错误和第二类错误。此外,建立了所提检验的渐近最优性,通过大量数值实验验证了其有效性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are rapidly gaining enormous popularity in recent years. However, the training of LLMs has raised significant privacy and legal concerns, particularly regarding the distillation and inclusion of copyrighted materials in their training data without proper attribution or licensing, an issue that falls under the broader concern of data misappropriation. In this article, we focus on a specific problem of data misappropriation detection, namely, to determine whether a given LLM has incorporated the data generated by another LLM. We propose embedding watermarks into the copyrighted training data and formulating the detection of data misappropriation as a hypothesis testing problem. We develop a general statistical testing framework, construct test statistics, determine optimal rejection thresholds, and explicitly control type I and type II errors. Furthermore, we establish the asymptotic optimality properties of the proposed tests, and demonstrate the empirical effectiveness through intensive numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。