用16个轻量级角色模型验证了大模型在博弈实验中的表现,发现粗略匹配人类数据即可通过验证。
Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel
- 固定16个角色配置,通过提示词控制行为,模拟人类决策。
- 仅一个条件未达标,差距仅0.011,整体表现符合预设标准。
- 适合关注大模型行为可控性与实验设计可靠性的研究者。
大型语言模型被越来越多地用作合成研究参与者,通常通过其边际响应是否与人类数据相似来验证。本文研究了一个由十六个轻量级角色化GPT-4.1配置组成的固定面板,在重复策略博弈中的表现。该面板在四个重复博弈单元中的三个达到了预注册的宽泛参考条件均值标准;唯一未达标项仅低于下限参考值0.011。变异强烈依赖提示词,但其占比受不确定性假设影响:固定面板对称狄利克雷敏感度下,提示间占比中位数为63%-71%(杰弗里斯α=0.5)和47%-53%(α=1),而有限机会插值估计为85%-96%。总体继续概率差异为+0.083和+0.078,保守联合95%置信区间分别为[-0.171, +0.330]和[-0.181, +0.330]。干预同时改变了继续过程及其文本表达。另一组改写与位置操作使合作率从0/40提升至37/40,裸配置下标签冲突亦揭示表达控制能力。原始人格级p13结果未进行前瞻性家族控制,事后精确门控结构上力不从心;因此p13应视为复现目标而非发现。外部评审暴露了家族误差、依赖性、构念及边界不确定性缺陷,零调用再分析改变解读但未重写历史记录。注册的边际标准可在不精确估计治疗-响应对象的情况下通过。公开胶囊验证了4,916次确认性第3-5阶段运行,无需实时模型调用。结果基于单一固定模型-提示面板,不证明人类可替代性。
原文摘要 · Abstract (English)
Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble human data. We study a fixed panel of sixteen lightweight persona-conditioned GPT-4.1 configurations in repeated strategic games. The panel met preregistered broad-reference condition-mean criteria in three of four repeated-game cells; the sole miss was 0.011 below the lower reference bound. Variation was strongly prompt-indexed, but its share depended on uncertainty assumptions: fixed-panel symmetric-Dirichlet sensitivities produced median between-prompt shares of 63%-71% under Jeffreys alpha=0.5 and 47%-53% under alpha=1, while finite-opportunity plug-in estimates were 85%-96%. Aggregate continuation-probability contrasts were +0.083 and +0.078, with conservative simultaneous 95% intervals [-0.171, +0.330] and [-0.181, +0.330]. The treatment jointly changed the continuation process and its textual representation. A separate wording-and-position operation shifted cooperation from 0/40 to 37/40 in the bare configuration, and a label conflict also revealed representation control. The original persona-level p13 result was not prospectively family-controlled, while a post-adjudication exact gate was structurally underpowered; p13 is therefore a replication target rather than a finding. External review exposed family-error, dependence, construct, and boundary-uncertainty defects, and zero-call reanalysis changed the interpretation without rewriting the historical record. The registered marginal criteria could be passed without precisely estimating the treatment-response object. A public capsule verifies 4,916 confirmatory Phase 3-5 runs with no live model calls. The results concern one fixed model-prompt panel and do not establish human substitutability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。