提出可抵抗重训练攻击的表格数据水印技术,确保版权可验证
RaMark: Radioactive Watermarking for Generated Tabular Data

- 将水印嵌入数据分布本身,使其成为生成模型不可剥离的组成部分
- 在十万级独立数据持有者场景下,显著优于七种主流方法
- 即使被重训练或篡改,水印仍可检测,适合隐私数据共享场景
生成模型的进步使合成表格数据成为敏感数据共享的可行方案,水印可用于所有权验证。然而,现有水印方法在重训练攻击下失效:攻击者用带水印数据重新训练模型,生成的新数据不再携带水印。为此,我们提出放射性水印机制(RaMark),通过将正弦依赖关系作为数据分布的内在成分嵌入,使水印与数据分布耦合。理论上证明,移除水印会显著降低数据效用并改变分布。在两个真实世界表格数据集上,大规模所有权验证实验(10^5个独立数据持有者)表明,RaMark在抵抗重训练和数据修改攻击方面显著优于七种先进方法。
原文摘要 · Abstract (English)
Recent advances in generative modeling have made generated tabular data a practical solution for privacy-sensitive data sharing, where watermarking enables ownership verification. However, existing watermarking methods fundamentally fail under retraining attacks, in which an adversary retrains a generative model on a watermarked dataset and regenerates high-utility data that no longer carries the watermark. We address this challenge by introducing radioactivity, the property that a watermark remains detectable after generative model retraining, and propose RaMark, a radioactive watermarking method that embeds a sinusoidal dependency as an intrinsic component of the data distribution. By coupling the watermark with the underlying distribution, RaMark ensures that any generative model preserving data utility also has to preserve the watermark. We theoretically show that with high probability removing watermark degrades utility and alters data distribution. Extensive experiments on two real-world tabular datasets, under a large-scale ownership verification setting with $10^5$ independent data owners, demonstrate that RaMark achieves substantially stronger radioactivity than seven state-of-the-art methods and consistently outperforms them against both retraining and data modification attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。