arXiv:2607.26348cs.CLcs.AI2026-07综述被引 1

LLM模拟用户回答常出错,导致决策偏差,需警惕其在调研中的应用风险。

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

论文配图:When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
图 1 · 摘自论文原文
  • 用统一协议测试多个LLM在真实调查数据上的表现
  • 所有模型在个体层面均不如非LLM基线,且高估身份对态度的影响
  • 结果适用于跨文化、多领域场景,提醒决策者慎用合成用户数据

大型语言模型(LLMs)正被广泛用作合成用户,模拟人类响应以支持产品、政策和市场决策。本文通过跨领域基准测试评估这种替代的有效性与局限性。采用统一协议,在两个独立真实数据集(美国社会态度调查GSS与世界价值观调查WVS)上,对四个覆盖8B到前沿能力的LLM模型进行评估。所有模型均对比了基于真实人类数据训练的非LLM基线。结果显示:第一,在个体层面,无一模型优于最强基线;在跨文化价值观问题上,模型表现显著低于基线,且该差距在距离感知和正确评分下仍存在。第二,模型系统性高估人口属性对态度的预测力,此偏差几乎出现在所有问题组合中,且对编码不变测量亦稳健。更大或更强大的模型也无法缓解此问题。决策影响分析显示:在细分群体任务中,模型使群体间差异放大2至4倍,导致美国半数和跨文化多数案例误判目标群体,并虚构不存在的真实群体差异。本文提供跨域基准与评估框架,供团队预先判断合成用户数据是否可用于决策支持。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model falls well below it, and the gap survives distance-aware and proper scoring. Second, models systematically over-determine demographics, treating identity as far more predictive of attitudes than it is among real people, a distortion present for nearly every question-group combination and robust to a coding-invariant measure. Neither failure is remedied by a larger, more capable model. A decision-impact analysis shows why this matters in practice: on a segment-targeting task the models inflate between-segment gaps two to fourfold, would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people. We make the cross-domain benchmark and the evaluation framework available on request, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.

大模型评估合成数据调研偏差决策风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。