用约束编程生成精准匹配统计数据的合成人群,无需个人数据。
Exact Synthetic Populations for Scalable Societal and Market Modeling
- 直接编码统计总量与结构关系,确保个体一致性。
- 生成的人群完全匹配目标人口分布,无抽样偏差。
- 适合政策模拟、市场分析等需可复现结果的场景。
我们提出一种基于约束编程的合成人口生成框架,可高精度重现目标统计数据并保证个体层面的一致性。与依赖样本推断分布的数据驱动方法不同,该方法直接编码聚合统计量与结构关系,实现对人口特征的精确控制,且无需使用微观数据。我们在官方人口数据源上验证了该方法,并研究了分布偏差对下游分析的影响。本工作隶属于 Emotia 公司的 Pollitics 项目,合成人口可通过大语言模型查询,用于建模社会行为、探索市场与政策情景,提供无需个人数据的可复现决策级洞察。
原文摘要 · Abstract (English)
We introduce a constraint-programming framework for generating synthetic populations that reproduce target statistics with high precision while enforcing full individual consistency. Unlike data-driven approaches that infer distributions from samples, our method directly encodes aggregated statistics and structural relations, enabling exact control of demographic profiles without requiring any microdata. We validate the approach on official demographic sources and study the impact of distributional deviations on downstream analyses. This work is conducted within the Pollitics project developed by Emotia, where synthetic populations can be queried through large language models to model societal behaviors, explore market and policy scenarios, and provide reproducible decision-grade insights without personal data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。