用真实人口数据让大模型更真实模拟灾害中人们的行为。
Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
- 将人口、时间使用和城市空间数据融入模型初始化与决策
- 灾害模拟中相关性从0.349提升至0.836,误差降低88%
- 适合做应急规划、城市韧性研究的学者和从业者
大语言模型(LLM)代理在缺乏历史先例的灾害与基础设施中断情境下,可生成人类行为模拟,但其个体合理性未必反映真实人群行为。本文评估了实证数据对提升模拟统计真实性的效果,构建了一个嵌入美国社区调查人口特征、美国时间使用调查基础作息及城市空间背景的实证接地式LLM代理框架,用于初始化、记忆、决策提示与活动执行。独立采集于2024年7月费城热浪期间的家庭调查作为外部验证基准。相较于无接地基线,接地模型在正常日常活动中,平均相关性由0.528提升至0.912,均方误差由0.066降至0.008;在热浪条件下,相关性由0.349升至0.836,均方误差由0.098降至0.012。接地模型捕捉到46.4%的观测热浪响应幅度,而基线仅20.6%。结果表明,实证接地能显著提升LLM代理在群体行为模拟中的统计可信度,同时揭示当前模型在人类适应机制建模上的不足。
原文摘要 · Abstract (English)
Large language model (LLM) agents offer a generative approach to simulating human behavior under conditions that may have few or no direct historical analogues, a common challenge in disaster and infrastructure-disruption planning. However, this generative capacity creates a validity problem: individually plausible agent reasoning may fail to reproduce empirical population behavior. We evaluate whether empirical grounding improves the statistical realism of LLM-agent simulations during disruptions. Specifically, we develop an empirically grounded LLM-agent framework that embeds demographic profiles from the American Community Survey, baseline routines from the American Time Use Survey, and urban spatial context into agent initialization, memory, decision prompts, and activity execution. An independent household survey conducted during the July 2024 Philadelphia heatwave is reserved as an external validation benchmark. Compared with an ungrounded LLM-agent baseline, the grounded model improved reconstruction of normal daily routines, increasing mean correlation with empirical activity profiles from 0.528 to 0.912 and reducing mean squared error from 0.066 to 0.008. Under heatwave conditions, the grounded model better reproduced survey-derived activity profiles, increasing mean correlation from 0.349 to 0.836 and reducing mean squared error from 0.098 to 0.012. The grounded model captured 46.4% of observed heatwave response amplitude, compared with 20.6% for the ungrounded baseline. These findings show that empirical grounding can make LLM agents more statistically credible simulators of population behavior while revealing remaining gaps in modeling human adaptation during disruptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。