比较四种生成数据方法在电网隐私与效用间的平衡表现。
Evaluating Privacy-Utility Tradeoffs in Synthetic Smart Grid Data
- 用四种生成模型合成家庭用电数据,对比效果。
- 扩散模型效用最高(宏平均F1达88.2%),CTGAN抗重构攻击最强。
- 适合关注隐私保护数据生成的能源系统研究者。
动态分时电价(dToU)的推广需要精准识别受益用户,但真实用电数据存在严重隐私风险,促使采用合成数据替代。本研究对比评估了四种合成数据生成方法:Wasserstein-GP生成对抗网络(WGAN)、条件表格式GAN(CTGAN)、扩散模型和高斯噪声增强,在不同合成条件下对分类效用、分布保真度和隐私泄露进行评估。结果表明,模型架构设计起关键作用:扩散模型实现最高效用(宏平均F1最高达88.2%),而CTGAN在抵御重构攻击方面表现最优。研究凸显结构化生成模型在构建隐私保护、数据驱动的能源系统中的潜力。
原文摘要 · Abstract (English)
The widespread adoption of dynamic Time-of-Use (dToU) electricity tariffs requires accurately identifying households that would benefit from such pricing structures. However, the use of real consumption data poses serious privacy concerns, motivating the adoption of synthetic alternatives. In this study, we conduct a comparative evaluation of four synthetic data generation methods, Wasserstein-GP Generative Adversarial Networks (WGAN), Conditional Tabular GAN (CTGAN), Diffusion Models, and Gaussian noise augmentation, under different synthetic regimes. We assess classification utility, distribution fidelity, and privacy leakage. Our results show that architectural design plays a key role: diffusion models achieve the highest utility (macro-F1 up to 88.2%), while CTGAN provide the strongest resistance to reconstruction attacks. These findings highlight the potential of structured generative models for developing privacy-preserving, data-driven energy systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。