arXiv:2411.10898cs.CRcs.LG2024-11被引 2

通过分布级修改,实现生成型分类数据的可验证水印。

Watermarking Generative Categorical Data

  • 将数据分布拆分为两部分,基于确定性关系嵌入密钥信号。
  • 利用逆解码后与原分布的总变差距离检测水印存在。
  • 适用于合成数据生成场景,适合关注分布而非具体数据点的研究者。

本文提出一种新颖的统计框架,用于生成型分类数据的水印。方法通过将数据分布分解为两个部分,基于确定性关系修改其中一个分布,从而在分布层面嵌入预设的密钥信号。验证时引入插入逆算法,通过测量逆解码数据与原始分布之间的总变差距离来检测水印。与以往主要针对具体数据集嵌入水印的方法不同,本方法在分布层面操作,支持从统计分布角度进行验证。这使其特别适合现代合成数据生成范式,其中数据分布本身比具体数据点更具重要性。方法的有效性通过理论分析和实证验证得到证明。

原文摘要 · Abstract (English)

In this paper, we propose a novel statistical framework for watermarking generative categorical data. Our method systematically embeds pre-agreed secret signals by splitting the data distribution into two components and modifying one distribution based on a deterministic relationship with the other, ensuring the watermark is embedded at the distribution-level. To verify the watermark, we introduce an insertion inverse algorithm and detect its presence by measuring the total variation distance between the inverse-decoded data and the original distribution. Unlike previous categorical watermarking methods, which primarily focus on embedding watermarks into a given dataset, our approach operates at the distribution-level, allowing for verification from a statistical distributional perspective. This makes it particularly well-suited for the modern paradigm of synthetic data generation, where the underlying data distribution, rather than specific data points, is of primary importance. The effectiveness of our method is demonstrated through both theoretical analysis and empirical validation.

数据水印生成模型统计方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。