arXiv:2607.28347cs.CLcs.SI2026-07

大模型难以自主模拟人类信念更新,需真实初始立场才有效。

LLMs struggle to simulate human belief updates in controlled environments

论文配图:LLMs struggle to simulate human belief updates in controlled environments
图 1 · 摘自论文原文
  • 用真人性格数据生成角色,让大模型模仿人类立场变化
  • 仅当输入真实初始立场时,部分模型能匹配人类最终立场分布
  • 大模型普遍存在中立倾向、小幅度偏移,且不会判断观点说服力

大模型被越来越多地用于社会科学研究中替代人类参与者,但其模拟效果尚未得到直接验证。我们测试了六种大模型在控制环境下模拟个体信念更新的能力,将模型输出与391名英国参与者的实际数据(来自Prolific平台)进行一对一比较,这些参与者在阅读Reddit评论后更新了对三个话题的立场。每位参与者由基于其人口统计学和人格特质数据构建的角色驱动的大模型模拟。结果发现,只有Qwen3-32B和GPT-5-Mini能在给定真实初始立场的情况下匹配人类的最终立场分布;所有模型均无法自动生成初始立场,也无法从自生立场出发产生真实的信念更新。三类系统性偏差普遍存在于所有模型中:中立立场占比过高、信念变动更频繁但幅度更小、无法正确评估评论的说服力。人格与人口统计特征角色对模拟精度无稳定影响。大模型对人类信念动态的模拟仅在具备真实起始条件时可靠,而当前多轮社交媒体模拟通常缺乏此类条件。

原文摘要 · Abstract (English)

LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants' actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.

大模型信念更新社会模拟认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。