arXiv:2510.09210cs.CRcs.LG2025-10NeurIPS被引 3

为无害数据投毒生成可验证的水印,防止滥用

Provable Watermarking for Data Poisoning Attacks

  • 在投毒后或同步时加入可证明的水印
  • 水印长度满足特定条件时能同时保证可检测与有效
  • 适合关注数据所有权和安全的从业者

近年来,数据投毒攻击逐渐设计得看似无害甚至有益,常用于验证数据集所有权或保护私有数据免遭非法使用。然而,这类发展可能引发误解与冲突,因数据投毒传统上被视为机器学习系统的安全威胁。为解决此问题,无害投毒生成者需能证明其生成数据的归属权,使用户可识别潜在投毒以防止误用。本文提出采用水印方案应对该挑战,引入两种可证明且实用的投毒水印方法:后投毒水印与投毒同步水印。分析表明,当后投毒水印长度为Θ(√d/ε_w),投毒同步水印长度在Θ(1/ε_w²)至O(√d/ε_p)之间时,水印化投毒数据可同时确保水印可检测性与投毒有效性,证实了水印在投毒攻击下的实用性。我们在多种攻击、模型与数据集上验证了理论结果。

原文摘要 · Abstract (English)

In recent years, data poisoning attacks have been increasingly designed to appear harmless and even beneficial, often with the intention of verifying dataset ownership or safeguarding private data from unauthorized use. However, these developments have the potential to cause misunderstandings and conflicts, as data poisoning has traditionally been regarded as a security threat to machine learning systems. To address this issue, it is imperative for harmless poisoning generators to claim ownership of their generated datasets, enabling users to identify potential poisoning to prevent misuse. In this paper, we propose the deployment of watermarking schemes as a solution to this challenge. We introduce two provable and practical watermarking approaches for data poisoning: {\em post-poisoning watermarking} and {\em poisoning-concurrent watermarking}. Our analyses demonstrate that when the watermarking length is $Θ(\sqrt{d}/ε_w)$ for post-poisoning watermarking, and falls within the range of $Θ(1/ε_w^2)$ to $O(\sqrt{d}/ε_p)$ for poisoning-concurrent watermarking, the watermarked poisoning dataset provably ensures both watermarking detectability and poisoning utility, certifying the practicality of watermarking under data poisoning attacks. We validate our theoretical findings through experiments on several attacks, models, and datasets.

数据投毒水印技术安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。