构建50万条模拟数据集,系统评测各类缺失值填补方法在社会经济数据中的表现。
IMAGIC-500: IMputation benchmark on A Generative Imaginary Country (500k samples)
- 基于世界银行合成数据生成50万样本的层级化模拟数据集
- 在10%~50%缺失率下对比多种方法对连续与分类变量的填补精度
- 适用于社会科学研究者评估填补算法性能,推动可复现研究
表格数据中的缺失值填补仍是数据科学与机器学习中的关键挑战,尤其在社会经济研究领域。然而,真实的社会经济数据常受严格隐私保护限制,难以公开共享,即使合成衍生数据也受限,严重制约了基准研究的可复现性与可及性。目前公开可用的合成数据集极为有限,导致针对社会经济数据的缺失值填补方法缺乏系统性评估基准。本研究利用世界银行公开的合成数据集《一个虚构国家的合成数据》,该数据集高度模拟真实世界银行家庭调查,且完全公开,便于方法研究。在此基础上,我们构建了IMAGIC-500数据集:选取约10万户家庭中的50万个人,包含19个社会经济特征,反映真实家庭调查的层级结构。本文在多种缺失机制(MCAR、MAR、MNAR)和缺失率(10%、20%、30%、40%、50%)下,建立全面的缺失值填补基准。评估涵盖连续与分类变量的填补精度、计算效率,以及对下游预测任务(如个体教育水平估计)的影响。结果揭示了统计方法、传统机器学习与深度学习(包括最新扩散模型)在填补效果上的优劣。IMAGIC-500数据集与基准旨在促进鲁棒填补算法的发展,支持可复现的社会科学研究。
原文摘要 · Abstract (English)
Missing data imputation in tabular datasets remains a pivotal challenge in data science and machine learning, particularly within socioeconomic research. However, real-world socioeconomic datasets are typically subject to strict data protection protocols, which often prohibit public sharing, even for synthetic derivatives. This severely limits the reproducibility and accessibility of benchmark studies in such settings. Further, there are very few publicly available synthetic datasets. Thus, there is limited availability of benchmarks for systematic evaluation of imputation methods on socioeconomic datasets, whether real or synthetic. In this study, we utilize the World Bank's publicly available synthetic dataset, Synthetic Data for an Imaginary Country, which closely mimics a real World Bank household survey while being fully public, enabling broad access for methodological research. With this as a starting point, we derived the IMAGIC-500 dataset: we select a subset of 500k individuals across approximately 100k households with 19 socioeconomic features, designed to reflect the hierarchical structure of real-world household surveys. This paper introduces a comprehensive missing data imputation benchmark on IMAGIC-500 under various missing mechanisms (MCAR, MAR, MNAR) and missingness ratios (10\%, 20\%, 30\%, 40\%, 50\%). Our evaluation considers the imputation accuracy for continuous and categorical variables, computational efficiency, and impact on downstream predictive tasks, such as estimating educational attainment at the individual level. The results highlight the strengths and weaknesses of statistical, traditional machine learning, and deep learning imputation techniques, including recent diffusion-based methods. The IMAGIC-500 dataset and benchmark aim to facilitate the development of robust imputation algorithms and foster reproducible social science research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。