用扩散模型生成更真实多样的合成人口,解决高维数据稀疏问题。
Generating Feasible and Diverse Synthetic Populations Using Diffusion Models
- 基于扩散模型建模人口属性联合分布,恢复缺失组合。
- 相比VAE和GAN,生成结果更符合实际且多样性更高。
- 适合交通仿真、城市规划等需要高质量合成数据的场景。
人口合成是生成具有现实特征的合成人群的关键任务,广泛应用于基于智能体的建模(ABM),尤其在交通系统分析中至关重要。当描述个体的属性维度较高时,调查数据难以支撑属性间的联合分布,导致数据稀疏,难以准确建模。深度生成模型虽可填补样本中不存在但实际存在的属性组合(采样零点),但易生成不合理的组合(结构零点)。本文提出一种基于扩散模型的人口合成方法,能有效恢复大量缺失的采样零点,同时最小化结构零点的生成。与变分自编码器(VAE)和生成对抗网络(GAN)相比,该方法在边缘分布相似性、可行性及多样性等多项指标上表现更优,实现了可行性与多样性的更好平衡。
原文摘要 · Abstract (English)
Population synthesis is a critical task that involves generating synthetic yet realistic representations of populations. It is a fundamental problem in agent-based modeling (ABM), which has become the standard to analyze intelligent transportation systems. The synthetic population serves as the primary input for ABM transportation simulation, with traveling agents represented by population members. However, when the number of attributes describing agents becomes large, survey data often cannot densely support the joint distribution of the attributes in the population due to the curse of dimensionality. This sparsity makes it difficult to accurately model and produce the population. Interestingly, deep generative models trained from available sample data can potentially synthesize possible attribute combinations that present in the actual population but do not exist in the sample data(called sampling zeros). Nevertheless, this comes at the cost of falsely generating the infeasible attribute combinations that do not exist in the population (called structural zeros). In this study, a novel diffusion model-based population synthesis method is proposed to estimate the underlying joint distribution of a population. This approach enables the recovery of numerous missing sampling zeros while keeping the generated structural zeros minimal. Our method is compared with other recently proposed approaches such as Variational Autoencoders (VAE) and Generative Adversarial Network (GAN) approaches, which have shown success in high dimensional tabular population synthesis. We assess the performance of the synthesized outputs using a range of metrics, including marginal distribution similarity, feasibility, and diversity. The results demonstrate that our proposed method outperforms previous approaches in achieving a better balance between the feasibility and diversity of the synthesized population.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。