arXiv:2607.10554stat.MEcs.LG2026-07

为表格数据设计可检测的水印方案,单样本即可识别。

Observation-Level Watermarking and Detection for Tabular Data

  • 基于观测级水印框架,兼容各类数据分布。
  • 即使仅有一个样本也能可靠检测水印。
  • 适合保护真实表格数据版权,抗子集攻击。

随着生成式AI的发展,水印技术被广泛用于检测AI生成数据的真实性并保护用户与创作者权益。尽管在图像和文本数据中已广泛应用,表格数据的水印研究仍不充分。现有方法主要针对数值型数据,对离散、类别及混合型数据关注较少。本文提出STAMP(Single-observation Tabular Attribution and Marking Procedure)——一种新型表格数据水印框架,能适应并保留广泛的数据分布特性。同时构建了相应的检测机制,可在样本量小至1时仍可靠识别水印。我们建立了渐近一致性和检测准确性的理论保证。通过大量模拟实验及两个真实数据应用,验证了该方法在子集操作下依然有效且保持数据保真度和高检测率。

原文摘要 · Abstract (English)

With the development of generative AI, watermarking techniques have been widely used to detect the authenticity of AI-generated data and protect the rights of users and creators. While it is already well applied in data types including imaging and text data, watermarking tabular data is still under-explored. Existing methods primarily focus on numerical data, leaving discrete, categorical, and mixed data less studied. In this work, we propose STAMP (Single-observation Tabular Attribution and Marking Procedure), a novel framework for watermarking tabular data that can accommodate and preserve a wide range of distributions. We also develop a corresponding detection mechanism, which can reliably identify watermarks even when the sample size is as small as one. We establish theoretical guarantees for asymptotic consistency and detection accuracy. Finally, through extensive simulation studies and two real-data applications, we demonstrate that the proposed method is effective and robust to subsetting, while maintaining data fidelity and a high detection rate.

数据水印表格数据版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。