arXiv:2510.10513cs.LG2025-10

用混合方法生成医疗表格数据,精准还原真实分布并保护隐私。

A Hybrid Machine Learning Approach for Synthetic Data Generation with Post Hoc Calibration for Clinical Tabular Datasets

  • 融合五种增强技术,用强化学习动态分配权重。
  • 生成数据与真实数据的统计差异极小,Wasserstein距离低至0.001。
  • 适合需要高保真合成数据的医疗AI研究者使用。

医疗研究因数据稀缺和严格的隐私法规(如HIPAA、GDPR)受限,难以获取真实医学数据。本文提出一种混合机器学习框架,用于生成高保真医疗表格数据,结合噪声注入、插值、高斯混合模型(GMM)采样、条件变分自编码器(CVAE)采样和SMOTE,并通过强化学习动态调节权重。关键创新在于多种校准技术:矩匹配、全直方图匹配、软及自适应软直方图匹配、迭代优化,有效对齐边缘分布并保留特征间联合依赖。在乳腺癌威斯康星(UCI库)和库尔纳医学院心脏病数据集上测试,生成数据的Wasserstein距离低至0.001,柯尔莫戈洛夫-斯米尔诺夫统计量约0.01,边际差异接近零;成对趋势得分超90%,最近邻对抗准确率接近50%,表明强隐私保护。基于合成数据训练的下游分类器准确率高达94%,F1分数超过93%,与真实数据训练模型相当。该方法可扩展、隐私友好,达到当前先进水平,为医疗敏感AI应用树立新基准。

原文摘要 · Abstract (English)

Healthcare research and development face significant obstacles due to data scarcity and stringent privacy regulations, such as HIPAA and the GDPR, restricting access to essential real-world medical data. These limitations impede innovation, delay robust AI model creation, and hinder advancements in patient-centered care. Synthetic data generation offers a transformative solution by producing artificial datasets that emulate real data statistics while safeguarding patient privacy. We introduce a novel hybrid framework for high-fidelity healthcare data synthesis integrating five augmentation methods: noise injection, interpolation, Gaussian Mixture Model (GMM) sampling, Conditional Variational Autoencoder (CVAE) sampling, and SMOTE, combined via a reinforcement learning-based dynamic weight selection mechanism. Its key innovations include advanced calibration techniques -- moment matching, full histogram matching, soft and adaptive soft histogram matching, and iterative refinement -- that align marginal distributions and preserve joint feature dependencies. Evaluated on the Breast Cancer Wisconsin (UCI Repository) and Khulna Medical College cardiology datasets, our calibrated hybrid achieves Wasserstein distances as low as 0.001 and Kolmogorov-Smirnov statistics around 0.01, demonstrating near-zero marginal discrepancy. Pairwise trend scores surpass 90%, and Nearest Neighbor Adversarial Accuracy approaches 50%, confirming robust privacy protection. Downstream classifiers trained on synthetic data achieve up to 94% accuracy and F1 scores above 93%, comparable to models trained on real data. This scalable, privacy-preserving approach matches state-of-the-art methods, sets new benchmarks for joint-distribution fidelity in healthcare, and supports sensitive AI applications.

合成数据医疗AI隐私保护数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。