arXiv:2603.27312cs.LG2026-03被引 1

用持续对比散度加速人口合成,突破传统方法规模瓶颈。

Scalable Maximum Entropy Population Synthesis via Persistent Contrastive Divergence

  • 引入持久对比散度,用随机采样替代全空间求和计算期望。
  • 在50个属性下仍保持0.018以内误差,运行时间仅随属性数线性增长。
  • 适合需要高多样性模拟的智慧城市建模,显著优于传统方法。

最大熵建模为从汇总普查数据生成合成人口提供了理论框架,无需个体微观数据。精确枚举方法的瓶颈在于对完整元组空间 $\cX$ 的显式求和计算期望,当类别属性数 $K \approx 20$ 时即不可行;现有采样方法依赖马尔可夫链蒙特卡洛,需调参与拒绝步骤。本文提出 extit{GibbsPCDSolver},基于持续对比散度(PCD)的随机替代方案:每次梯度步使用 $N$ 个合成个体组成的持久池,通过吉布斯扫掠更新,无需显式构造 $\cX$ 即可提供模型期望的随机近似。在可控基准与 $K=15$ 的意大利人口基准 extit{Syn-ISTAT} 上验证,该方法在 $K \in \{12, 20, 30, 40, 50\}$ 下均保持 $\MRE \in [0.010, 0.018]$,而 $|\cX|$ 增长十八个数量级,运行时间仅为 $O(K)$ 而非 $O(|\cX|)$。在 extit{Syn-ISTAT} 上,模型训练约束下达到 $\MRE=0.03$,且生成人口有效样本量 $\Neff = N$,远超广义重加权法的 $\Neff \approx 0.012\,N$,多样性提升86.8倍,对城市仿真至关重要。

原文摘要 · Abstract (English)

Maximum entropy (MaxEnt) modelling provides a principled framework for generating synthetic populations from aggregate census data, without access to individual-level microdata. The bottleneck of exact-enumeration approaches is expectation computation by explicit summation over the full tuple space $\cX$, which becomes infeasible for more than $K \approx 20$ categorical attributes; sampling-based alternatives exist but rely on Metropolis-type schemes that require proposal tuning and rejection steps. We propose \emph{GibbsPCDSolver}, a stochastic replacement for this computation based on Persistent Contrastive Divergence (PCD): a persistent pool of $N$ synthetic individuals is updated by Gibbs sweeps at each gradient step, providing a stochastic approximation of the model expectations without ever materialising $\cX$. We validate the approach on controlled benchmarks and on \emph{Syn-ISTAT}, a $K{=}15$ Italian demographic benchmark with analytically exact marginal targets derived from ISTAT-inspired conditional probability tables. Scaling experiments across $K \in \{12, 20, 30, 40, 50\}$ confirm that GibbsPCDSolver maintains $\MRE \in [0.010, 0.018]$ while $|\cX|$ grows eighteen orders of magnitude, with runtime scaling as $O(K)$ rather than $O(|\cX|)$. On Syn-ISTAT, GibbsPCDSolver reaches $\MRE{=}0.03$ on training constraints and -- crucially -- produces populations with effective sample size $\Neff = N$ versus $\Neff \approx 0.012\,N$ for generalised raking, an $86.8{\times}$ diversity advantage that is essential for agent-based urban simulations.

人口合成最大熵对比散度城市仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。