arXiv:2604.15329cs.HCcs.AI2026-04被引 1

用大模型模拟人类行为实验,发现能复现部分趋势但效果量不一致。

Evaluating LLMs as Human Surrogates in Controlled Experiments

  • 将人类数据转为结构化提示,用现成大模型生成响应
  • 大模型复现了人类在准确性感知中的多个方向性效应
  • 适合想低成本验证实验设计的研究者使用

大型语言模型(LLMs)被越来越多用于行为研究中模拟人类反应,但其生成数据是否能支持与人类数据相同的实验推断仍不明确。我们通过直接比较现成大模型生成的回应与真实人类在经典准确性感知调查实验中的回应来评估这一点。每个真实人类观测值被转化为结构化提示,模型在无任务特定训练的情况下生成0-10的单一结果变量;对人类与合成数据采用相同的统计分析方法。结果发现,大模型再现了人类观察到的多个方向性效应,但效应大小和调节模式在不同模型间存在差异。因此,现成大模型在受控条件下可捕捉群体信念更新模式,但并不总是匹配人类尺度的效应,明确了大模型生成数据作为行为代理的适用边界。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to simulate human responses in behavioral research, yet it remains unclear when LLM-generated data support the same experimental inferences as human data. We evaluate this by directly comparing off-the-shelf LLM-generated responses with human responses from a canonical survey experiment on accuracy perception. Each human observation is converted into a structured prompt, and models generate a single 0--10 outcome variable without task-specific training; identical statistical analyses are applied to human and synthetic responses. We find that LLMs reproduce several directional effects observed in humans, but effect magnitudes and moderation patterns vary across models. Off-the-shelf LLMs therefore capture aggregate belief-updating patterns under controlled conditions but do not consistently match human-scale effects, clarifying when LLM-generated data can function as behavioral surrogates.

大模型行为研究实验模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。