arXiv:2607.18310physics.soc-phcs.AI2026-07

用真实数据发现大模型个体模拟会坍缩分布,提出先建分布再分配的修正方案。

Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling

  • 先建群体分布再分配角色,避免独立建模导致的分布坍缩
  • 用口语化采样修复分布但过度分散,需配合预算路由控制
  • 适合做社会模拟、政策推演的研究者,尤其关注非西方人群

将每个个体视为独立大语言模型代理的合成人口工具存在根本缺陷。基于2,414名真实世界价值观调查受访者数据,在非西方(土耳其优先)语境下,我们发现:独立建模的2,414个代理在四种情景下五次种子测试中,响应分布集中度从0.36升至0.69,熵从1.46降至0.77,85%出现坍缩,总变差距离(TVD)达0.44;其坍缩程度与题目结构相关(相关系数r=0.55)。口语化采样(VS)在三个模型族中无需训练即提升问卷拟合度(+7到+10),但在Qwen上显著(p=0.002,d=6.2),却普遍引发过分散(标准差比0.4-0.56升至1.26-1.37)。问卷拟合度仅弱关联于代理行为:在单一任务中,角色被最低价默认主导(约80%),收入仅调节舒适选择(0%→7%→32%)。置换对照记忆攻击与选举回测显示,尽管总体强度保留,但子群和个体主张受回忆偏差与不确定性污染。我们提出修正:一次建模分布(使用VS),以近似常数成本分配给具体角色,搭配预算感知路由器(诚实AUC为0.805,而非代码源的1.0伪真值)。核心贡献不依赖现实性宣称,而是揭示独立代理路径的内生不一致及分布优先路径的校准条件。

原文摘要 · Abstract (English)

Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent. Using real survey microdata, we show that this paradigm has a basic failure mode, and we set a distribution-first corrective against it, all measured with a deterministic, construct-validated verifier on non-WEIRD (Turkey-first) data. First, N independent LLM agents grounded on 2,414 real World Values Survey respondents fail to reproduce the population's response distribution: they pile onto a modal default (four scenarios x five seeds: concentration 0.36->0.69, entropy 1.46->0.77, 85% collapse, TVD=0.44), and the collapse is a predictable function of scenario structure (r=0.55 with a single-answer structure). Second, Verbalized Sampling (VS) fixes the field's chronic under-dispersion without training in three model families (fidelity +7 to +10; significant on Qwen, p=0.002, d=6.2), yet the same move universally overshoots into over-dispersion (SD-ratio 0.4-0.56 -> 1.26-1.37), a structural property of VS. Third, survey fidelity transfers only weakly to agentic behavior: in a single-model, single-domain booking task, a persona is dominated by a cheapest-default (~80%) that income modulates but does not override (comfort choice 0%->7%->32% across income bands). Fourth, a placebo-controlled memorization attack and an election backtest show VS keeps aggregate strength while subgroup and individual claims are contaminated by recall and underdetermination. We close with the corrective: model the distribution once (VS) and assign it to grounded characters at O(1) cost, with a budget-aware router whose honest AUC is 0.805, not the tautological 1.0 of a code-derived oracle. The central contribution needs no realism claim: it measures the internal inconsistency of the independent-agent route and the conditions under which the distribution-first route calibrates.

社会模拟分布建模大模型代理非西方数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。