arXiv:2502.17773stat.MEcs.AI2025-02综述被引 10

用不确定性量化方法评估大模型模拟调查数据的等效人类样本量。

How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective

  • 通过自适应选择模拟数量,生成有保证的置信区间。
  • 不同大模型在各领域模拟效果差异明显,有效人类样本量不一。
  • 提供可量化的大模型仿真质量指标,适合政策与社会研究者参考。

大语言模型(LLMs)被越来越多地用于模拟调查回应,但合成数据可能与真实人群存在偏差,导致推断不可靠。本文提出一个通用框架,将LLM生成的回应转化为对人类回应总体参数的可靠置信集,量化由人-模型偏差引发的不确定性。关键在于模拟样本量的选择:过多导致置信集过窄且覆盖不足,过少则使置信集过宽并受随机噪声主导。我们提出一种数据驱动的方法,自适应选择样本量以实现名义平均覆盖率,不受模型仿真精度或置信集构建方式影响。所选样本量还反映了该模型能代表的有效人类群体规模,提供了仿真精度的定量度量。在真实调查数据集上的实验显示,不同大模型在各领域的仿真精度存在显著差异。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to simulate survey responses, but synthetic data can be misaligned with the human population, leading to unreliable inference. We develop a general framework that converts LLM-simulated responses into reliable confidence sets for population parameters of human responses, quantifying the uncertainty induced by the human-LLM misalignment. The key design choice is the number of simulated responses: too many produce overly narrow sets with poor coverage, while too few yield overly wide and uninformative sets dominated by stochastic noise. We propose a data-driven approach that adaptively selects the simulation sample size to achieve nominal average-case coverage, regardless of the LLM's simulation fidelity or the confidence set construction procedure. The selected sample size is further shown to reflect the effective human population size that the LLM can represent, providing a quantitative measure of its simulation fidelity. Experiments on real survey datasets reveal heterogeneous simulation fidelity across different LLMs and domains.

大模型评估不确定性量化调查模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。