用约束优化生成医疗表格数据,兼顾隐私与临床价值。
PSyGenTAB: A Privacy-Preserving Framework for Synthetic Clinical Tabular Data Generation via Constrained Optimization
- 将生成过程建模为带隐私约束的优化问题,直接控制隐私风险。
- 合成数据保留关键临床关系和少数类诊断模式,模型性能接近真实数据。
- 适合需要跨机构共享数据的医疗AI研究者使用。
医疗AI发展受限于高质量临床数据的获取困难,受机构数据孤岛和严格隐私法规(如HIPAA、GDPR)制约。现有合成数据方法缺乏对隐私-效用权衡的显式管理,常导致临床模式丢失或患者重识别风险。我们提出PSyGenTAB,一种基于约束优化的隐私保护生成框架,采用增广拉格朗日法求解,将可配置隐私约束嵌入训练过程,确保最低隐私阈值的同时最大化数据效用。在多个临床基准测试中,该方法有效保留特征间临床关联及少数类别诊断模式。下游评估显示,使用合成数据训练的模型在‘以合成数据训练,真实数据测试’和‘以真实数据训练,合成数据测试’两种协议下表现与真实数据训练模型相当。隐私审计表明,精确记录重现率降低,且对成员推断攻击具有强鲁棒性。结果证明PSyGenTAB是平衡隐私保护与临床效用的可靠框架,支持安全的跨机构医疗AI开发。
原文摘要 · Abstract (English)
The development of medical AI is constrained by limited access to high-quality clinical data due to institutional silos and strict privacy regulations such as HIPAA and GDPR. Synthetic data generation offers a potential solution, but existing methods lack principled mechanisms to explicitly manage the privacy-utility trade-off, often degrading clinically meaningful patterns or risking patient re-identification. We present PSyGenTAB, a privacy-preserving generative framework that formulates synthetic healthcare data generation as a constrained optimization problem solved using the Augmented Lagrangian Method. By embedding configurable privacy constraints directly into model training, PSyGenTAB enforces minimum privacy thresholds while maximizing clinical data utility. Across multiple clinically motivated benchmarks, PSyGenTAB preserves inter-feature clinical relationships and minority-class diagnostic patterns essential for reliable health AI. Downstream evaluation using Train-on-Synthetic, Test-on-Real and Train-on-Real, Test-on-Synthetic protocols shows that models trained on synthetic data achieve performance comparable to those trained on real patient records. Privacy auditing further demonstrates reduced exact record reproduction and strong resilience to membership inference attacks. These results establish PSyGenTAB as a principled framework for balancing privacy protection and clinical utility in synthetic healthcare data, supporting secure cross-institutional AI development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。