用生成模型联合融合多源数据,提升城市人口模拟的多样性与可行性
Enhancing Diversity and Feasibility: Joint Population Synthesis from Multi-source Data Using Generative Models
- 采用WGAN框架联合处理多源数据,统一建模复杂特征关系
- 相比传统方法,召回率提升7%,精确率提升15%,多样性显著增强
- 适合交通规划、城市模拟等需要高保真人口数据的研究者使用
生成真实合成人口对交通与城市规划中的基于代理的模型(ABM)至关重要。现有方法存在两大缺陷:一是依赖单一数据集或分步融合生成,难以捕捉特征间复杂关联;二是难以处理采样零(未观测但合理)与结构零(因逻辑约束不可行)组合,导致数据多样性和可行性不足。本文提出一种基于带梯度惩罚的Wasserstein生成对抗网络(WGAN-GP)的联合学习方法,通过引入反向梯度惩罚项作为生成器损失的正则化项,同时提升数据多样性和可行性。评估采用统一相似性指标,并重点以召回率、精确率和F1分数衡量多样性与可行性。结果表明,所提联合方法优于序列基线,召回率提高7%,精确率提升15%;正则化项进一步使召回率增加10%,精确率提升1%。综合五项指标评分,联合方法达88.1,高于基线的84.6。由于合成人口是ABM的关键输入,该多源生成方法有望显著提升ABM的准确性和可靠性。
原文摘要 · Abstract (English)
Generating realistic synthetic populations is essential for agent-based models (ABM) in transportation and urban planning. Current methods face two major limitations. First, many rely on a single dataset or follow a sequential data fusion and generation process, which means they fail to capture the complex interplay between features. Second, these approaches struggle with sampling zeros (valid but unobserved attribute combinations) and structural zeros (infeasible combinations due to logical constraints), which reduce the diversity and feasibility of the generated data. This study proposes a novel method to simultaneously integrate and synthesize multi-source datasets using a Wasserstein Generative Adversarial Network (WGAN) with gradient penalty. This joint learning method improves both the diversity and feasibility of synthetic data by defining a regularization term (inverse gradient penalty) for the generator loss function. For the evaluation, we implement a unified evaluation metric for similarity, and place special emphasis on measuring diversity and feasibility through recall, precision, and the F1 score. Results show that the proposed joint approach outperforms the sequential baseline, with recall increasing by 7\% and precision by 15\%. Additionally, the regularization term further improves diversity and feasibility, reflected in a 10\% increase in recall and 1\% in precision. We assess similarity distributions using a five-metric score. The joint approach performs better overall, and reaches a score of 88.1 compared to 84.6 for the sequential method. Since synthetic populations serve as a key input for ABM, this multi-source generative approach has the potential to significantly enhance the accuracy and reliability of ABM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。