arXiv:2506.06407cs.CRcs.AI2025-06NeurIPS被引 5

为时间序列扩散模型设计了首个直接在数据空间嵌入水印的方法。

TimeWak: Temporal Chained-Hashing Watermark for Time Series Data

  • 直接在时序特征数据空间中嵌入链式哈希水印,处理时空异质性与依赖关系。
  • 在五组数据上实现61.96%的上下文FID提升和8.44%的相关性得分提升。
  • 支持ε-精确反演,确保水印检测鲁棒性,适合隐私敏感数据共享场景。

由扩散模型生成的合成时间序列可帮助共享如患者功能磁共振记录等隐私敏感数据。合成数据的关键标准包括高数据效用和可追溯性以验证数据来源。现有水印方法多嵌入同质潜在空间,但主流时间序列生成器在数据空间运行,导致基于潜在空间的水印不兼容。这带来在数据空间直接水印的挑战,需应对特征异质性和时序依赖。本文提出TimeWak,首个面向多变量时间序列扩散模型的水印算法。为处理时序依赖与空间异质性,TimeWak直接在时序-特征数据空间嵌入时序链式哈希水印。另一独特特性是ε-精确反演,解决扩散过程反演中不同特征重建误差分布不均的问题。我们推导了多变量时间序列反演的误差边界,同时保证水印可检测性。我们在五组数据和多种时长基线上广泛评估TimeWak对合成数据质量、水印可检测性及抗后编辑攻击的鲁棒性。结果表明,TimeWak相较最强基线在上下文FID上提升61.96%,相关性得分提升8.44%,且保持持续可检测。

原文摘要 · Abstract (English)

Synthetic time series generated by diffusion models enable sharing privacy-sensitive datasets, such as patients' functional MRI records. Key criteria for synthetic data include high data utility and traceability to verify the data source. Recent watermarking methods embed in homogeneous latent spaces, but state-of-the-art time series generators operate in data space, making latent-based watermarking incompatible. This creates the challenge of watermarking directly in data space while handling feature heterogeneity and temporal dependencies. We propose TimeWak, the first watermarking algorithm for multivariate time series diffusion models. To handle temporal dependence and spatial heterogeneity, TimeWak embeds a temporal chained-hashing watermark directly within the temporal-feature data space. The other unique feature is the $ε$-exact inversion, which addresses the non-uniform reconstruction error distribution across features from inverting the diffusion process to detect watermarks. We derive the error bound of inverting multivariate time series while preserving robust watermark detectability. We extensively evaluate TimeWak on its impact on synthetic data quality, watermark detectability, and robustness under various post-editing attacks, against five datasets and baselines of different temporal lengths. Our results show that TimeWak achieves improvements of 61.96% in context-FID score, and 8.44% in correlational scores against the strongest state-of-the-art baseline, while remaining consistently detectable.

时间序列水印扩散模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。