用大模型模拟人群对政策的反应,提升预测准确性与可解释性。
LLM Powered Social Digital Twins: A Framework for Simulating Population Behavioral Response to Policy Interventions
- 以大模型作为个体认知引擎,构建虚拟人群数字孪生
- 在疫情案例中预测误差比传统方法低20.7%
- 适用于交通、经济、环保等多领域政策仿真
预测人群对政策干预的响应是计算社会科学与公共政策的核心挑战。传统方法依赖历史统计关联,缺乏机制解释力且难以应对新政策场景。本文提出一种通用框架,构建社会数字孪生——通过大语言模型(LLM)作为个体代理的认知引擎,每个代理基于人口与心理特征接收政策信号,输出多维度行为概率向量。校准层将代理响应聚合为可观测的宏观指标,实现与真实数据的验证,并支持反事实政策分析。我们在疫情响应领域进行实例化,以新冠为案例,利用丰富观测数据。在预留测试期内,校准后的数字孪生在六个行为类别上,宏平均预测误差较梯度提升基线降低20.7%。反事实实验显示响应具有单调性和边界性,具备行为合理性。该框架具有领域无关性,可应用于交通政策、经济干预、环境法规等任何影响群体行为的场景。我们讨论了其对政策模拟的意义、现有局限及扩展方向。
原文摘要 · Abstract (English)
Predicting how populations respond to policy interventions is a fundamental challenge in computational social science and public policy. Traditional approaches rely on aggregate statistical models that capture historical correlations but lack mechanistic interpretability and struggle with novel policy scenarios. We present a general framework for constructing Social Digital Twins - virtual population replicas where Large Language Models (LLMs) serve as cognitive engines for individual agents. Each agent, characterized by demographic and psychographic attributes, receives policy signals and outputs multi-dimensional behavioral probability vectors. A calibration layer maps aggregated agent responses to observable population-level metrics, enabling validation against real-world data and deployment for counterfactual policy analysis. We instantiate this framework in the domain of pandemic response, using COVID-19 as a case study with rich observational data. On a held-out test period, our calibrated digital twin achieves a 20.7% improvement in macro-averaged prediction error over gradient boosting baselines across six behavioral categories. Counterfactual experiments demonstrate monotonic and bounded responses to policy variations, establishing behavioral plausibility. The framework is domain-agnostic: the same architecture applies to transportation policy, economic interventions, environmental regulations, or any setting where policy affects population behavior. We discuss implications for policy simulation, limitations of the approach, and directions for extending LLM-based digital twins beyond pandemic response.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。