用合成数据预测肯尼亚牧民儿童疫苗漏打风险,既准确又保护隐私。
Using Synthetic Data for Machine Learning-based Childhood Vaccination Prediction in Narok, Kenya

- 用机器学习分析8年真实疫苗记录,识别高风险儿童。
- 合成数据训练模型,预测准确率超90%且不泄露隐私。
- 适合资源匮乏地区,兼顾隐私与精准干预的实用方案。
背景:低资源环境下数据利用不足阻碍了疫苗配送体系,尤其在游牧群体中,儿童易错过关键疫苗接种。以肯尼亚纳罗克郡马赛人为例,高质量、大规模数据的缺失导致免疫覆盖率估算不准,资源配置效率低下,难以及时干预。此外,敏感健康数据在脆弱群体中的隐私风险更高。目标:一是识别大规模人口中可能漏打疫苗的儿童,实现及时、基于证据的干预以提升接种率;二是更好地保护弱势群体的敏感健康数据隐私。方法:我们数字化了8年的国家卫生部510表(MOH 510)疫苗记录(n=6,913),并应用逻辑回归和XGBoost等机器学习模型进行风险预测。同时,采用一种新型基于表格扩散的合成数据生成方法(TabSyn)来保护模型中的患者隐私。结果:分类技术可可靠预测疫苗漏打风险,部分疫苗的召回率、精确率和F1分数均超过90%。使用合成数据训练模型而非真实数据,在不牺牲预测性能的前提下,有效保护了原始数据中个体的隐私。结论:这些结果支持在数字基础设施薄弱的诊所中引入合成数据策略,实现隐私保护、可扩展的儿童免疫覆盖率预测,推动健康信息系统的可持续发展。
原文摘要 · Abstract (English)
Background: Limited data utilization in low-resource settings poses a barrier to the vaccine delivery ecosystem, undermining efforts to achieve equitable immunization coverage. In nomadic populations, individuals face an increased risk of missing crucial vaccination doses as children. One such population is the Maasai in Narok County, Kenya, where the absence of high-volume, quality data hampers accurate coverage estimates, impedes efficient resource allocation, and weakens the ability to deliver timely interventions. Additionally, data privacy concerns are heightened in groups with limited sensitive data. Objectives: First, we aim to identify children at risk of missing key vaccines across a large population to provide timely, evidence-based interventions that support increased vaccination coverage. Second, we aim to better protect the privacy of sensitive health data in a vulnerable population. Methods: We digitized 8 years of child vaccination records from the MOH 510 registry (n=6,913) and applied machine learning models (Logistic Regression and XGBoost) to identify children at risk. Additionally, we utilize a novel approach to tabular diffusion-based synthetic data generation (TabSyn) to protect patient privacy within the models. Results: Our findings show that classification techniques can reliably and successfully predict children at risk of missing a vaccine, with recall, precision, and F1-scores exceeding 90% for some vaccines modeled. Additionally, training these models with synthetic data rather than real data, thus preserving the privacy of individuals within the original dataset, does not lead to a loss in predictive performance. Conclusion: These results support the use of synthetic data implementation in health informatics strategies for clinics with limited digital infrastructure, enabling privacy-preserving, scalable forecasting for childhood immunization coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。