用缺失数据训练生成模型,合成出接近真实人口的数据。
Population Synthesis using Incomplete Information
- 用掩码矩阵表示缺失值,改进WGAN使其能处理不完整个体数据。
- 在瑞典出行调查数据上验证,合成数据与完整数据训练结果高度相似。
- 适合隐私受限或数据缺失场景,推动生成模型在人口合成中的应用。
本文提出一种基于水军生成对抗网络(WGAN)的人口合成模型,用于在存在缺失信息的微观样本上进行训练。通过引入掩码矩阵表示缺失值,设计了一种适用于不完整数据的WGAN训练算法,以应对因隐私保护或数据采集限制导致的属性缺失问题。研究对比了在不完整微观样本和完整微观样本上训练的WGAN模型所生成的合成人口。基于瑞典全国出行调查数据进行多轮评估,通过将各模型生成的合成人口与真实人口数据集进行比较,验证了所提方法的有效性。实验结果显示,该方法生成的合成数据在统计特性上与使用完整数据训练的模型及真实数据高度一致。本研究为不完整数据下的人口合成提供了稳健解决方案,拓展了深度生成模型在该领域的应用前景。
原文摘要 · Abstract (English)
This paper presents a population synthesis model that utilizes the Wasserstein Generative-Adversarial Network (WGAN) for training on incomplete microsamples. By using a mask matrix to represent missing values, the study proposes a WGAN training algorithm that lets the model learn from a training dataset that has some missing information. The proposed method aims to address the challenge of missing information in microsamples on one or more attributes due to privacy concerns or data collection constraints. The paper contrasts WGAN models trained on incomplete microsamples with those trained on complete microsamples, creating a synthetic population. We conducted a series of evaluations of the proposed method using a Swedish national travel survey. We validate the efficacy of the proposed method by generating synthetic populations from all the models and comparing them to the actual population dataset. The results from the experiments showed that the proposed methodology successfully generates synthetic data that closely resembles a model trained with complete data as well as the actual population. The paper contributes to the field by providing a robust solution for population synthesis with incomplete data, opening avenues for future research, and highlighting the potential of deep generative models in advancing population synthesis capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。