arXiv:2608.10891cs.LG2026-08

对比合成数据生成与去噪匿名化在隐私预测中的表现,发现简单方法更优。

Benchmarking Time Series Generation Methods for Privacy-Preserving Forecasting

论文配图:Benchmarking Time Series Generation Methods for Privacy-Preserving Forecasting
图 1 · 摘自论文原文
  • 用合成数据训练、真实数据测试评估生成方法性能
  • 噪声匿名化最安全但预测效果最差,简单转换方法反超深度模型
  • 新模型Grasynda-P兼顾预测与隐私,位于最优权衡线上

在隐私敏感领域,时间序列预测常需基于公开数据而非原始观测值训练模型。合成时间序列生成主要用于数据增强,补充原始训练集;但当其完全替代原始数据时的表现及释放数据带来的隐私风险仍缺乏研究。本文通过一个‘用合成数据训练,真实数据测试’(TSTR)的基准评测,评估合成生成方法与基于噪声的匿名化基线在七组数据集上的表现,联合衡量预测性能与基于距离的隐私风险,刻画两者间的权衡关系。我们提出Grasynda-P,是图生成器Grasynda的隐私导向扩展,引入矩阵集成与核密度估计。结果表明:(1) 无生成方法能完全替代原始数据;(2) 噪声匿名化提供最强隐私保护但预测性能最差;(3) 在该设定下,基于简单变换的生成器优于深度生成模型;(4) Grasynda-P位于帕累托最优前沿,预测性能有竞争力且隐私分离更强于其他生成器。本基准为新型隐私感知合成时间序列生成方法的评估与发展提供了参考点。

原文摘要 · Abstract (English)

Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. Synthetic time series generation has been developed primarily for data augmentation, where generated series supplement the original training set. How well these methods perform when fully replacing the original data - and how much privacy risk the released series carry - remains underexplored. We address this gap through a benchmark evaluating synthetic generation methods and noise-based anonymization baselines under a Train on Synthetic, Test on Real (TSTR) protocol. We jointly assess forecasting performance and distance-based empirical privacy risk across seven datasets, characterizing the trade-off between these objectives. We also introduce Grasynda-P, a privacy-motivated extension of the graph-based generator Grasynda, incorporating matrix ensembling and kernel density estimation. Our results show that: (1) no generation method fully substitutes for original training data; (2) noise-based anonymization yields the strongest privacy but the worst forecasting performance; (3) simple transformation-based generators outperform deep generative models for forecasting in this setting; and (4) Grasynda-P lies on the Pareto frontier, achieving competitive forecasting with stronger privacy separation than other generators. This benchmark establishes a reference point for evaluating and developing new privacy-aware synthetic time series generation methods.

时间序列隐私生成合成数据评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。