arXiv:2511.05289cs.LG2025-11中稿 · as a proceedings p…

用生成数据增强防范医疗时间序列模型的成员推理攻击

Embedding-Space Data Augmentation to Prevent Membership Inference Attacks in Clinical Time Series Forecasting

  • 通过生成合成数据重训模型,混淆攻击者判断
  • ZOO-PCA方法使攻击者真假阳性率比下降最多
  • 适合关注医疗数据隐私的模型开发者

在电子健康记录(EHR)的时间序列预测任务中,如何在保障强隐私的同时维持高预测性能至关重要。本文研究数据增强对时间序列预测模型成员推理攻击(MIA)的缓解作用。结果表明,使用合成数据重新训练可显著降低基于损失的MIA有效性,从而减少攻击者的真正例率与假正例率之比。核心挑战在于生成既贴近原始训练数据、又能引入足够新异性的合成样本,以提升模型泛化能力并迷惑攻击者。我们评估了多种增强策略:零阶优化(ZOO)、受主成分分析约束的ZOO变体(ZOO-PCA)以及MixUp。实验显示,ZOO-PCA在不牺牲测试性能的前提下,对MIA的真正例率/假正例率比改善最为显著。

原文摘要 · Abstract (English)

Balancing strong privacy guarantees with high predictive performance is critical for time series forecasting (TSF) tasks involving Electronic Health Records (EHR). In this study, we explore how data augmentation can mitigate Membership Inference Attacks (MIA) on TSF models. We show that retraining with synthetic data can substantially reduce the effectiveness of loss-based MIAs by reducing the attacker's true-positive to false-positive ratio. The key challenge is generating synthetic samples that closely resemble the original training data to confuse the attacker, while also introducing enough novelty to enhance the model's ability to generalize to unseen data. We examine multiple augmentation strategies - Zeroth-Order Optimization (ZOO), a variant of ZOO constrained by Principal Component Analysis (ZOO-PCA), and MixUp - to strengthen model resilience without sacrificing accuracy. Our experimental results show that ZOO-PCA yields the best reductions in TPR/FPR ratio for MIA attacks without sacrificing performance on test data.

隐私保护医疗时序数据增强成员推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。