arXiv:2606.09013cs.CL2026-06综述

评估大模型在分布层面复现人类调查数据的能力,发现均值匹配不等于真实分布复现。

Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level

论文配图:Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level
图 1 · 摘自论文原文
  • 从分布角度评估大模型对人类调查数据的复现能力,涵盖二值、分类和计数三类变量。
  • 大模型虽能复现条件下的行为模式,但对购买数量分布的复现远差于简单基准模型。
  • 结构化角色与多模态输入提升复现效果,显式推理提示反而导致性能下降。

大模型被广泛用于模拟人类调查回答,但现有评估多基于均值或聚合一致性,难以揭示模型是否复现了人类行为的变异特征。本文基于一项非公开的2010年韩国方便面购买消费者选择实验数据,评估大模型在分布层面的复现表现。研究包含三类不同统计类型变量:二值购买发生、分类品牌选择、计数购买数量。针对每类变量,比较人类与大模型在均值、模式和分布层面的一致性,并与仅来自人类数据的参考基线对比。结果表明,大模型能较好复现条件级行为模式,但在分布结构上表现不佳:对于购买数量,无一模型优于不依赖条件的基线——该基线直接匹配全体人类数据的联合分布。值得注意的是,尽管某些模型在均值上表现良好,其分布距离人类仍可能高于该基线,说明仅用均值评估会误导结论。此外,输入配置影响复现效果:结构化人物设定和多模态输入提升对齐度,而显式推理提示则导致性能单调下降。

原文摘要 · Abstract (English)

LLMs are increasingly used to simulate human survey responses, but prior work has mainly evaluated replication using mean-level or aggregate agreement, offering limited insight into whether LLMs reproduce the variability of human behavior. We evaluate LLM-based survey replication at the distributional level using a non-public 2010 consumer choice experiment on Korean instant noodle purchases, a setting unlikely to overlap with model training data. We evaluate three response variables of differing statistical type: binary purchase incidence, categorical brand choice, and count purchase quantity. For each, we compare human and LLM responses at mean-level, pattern, and distributional alignment, and against reference baselines from the human data alone. LLMs reproduce condition-level patterns reasonably well but fail to capture distributional structure: for purchase quantity, no model beats a condition-insensitive baseline that simply matches the pooled human distribution. Because models that match human means well can still produce distributions further from humans than this baseline, mean-based evaluation alone can be actively misleading. Replication also varies with input configuration, with structured personas and multimodal inputs improving alignment while explicit reasoning prompting degrades it monotonically.

大模型评估分布复现调查模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。