arXiv:2606.30372cs.AIcs.CY2026-06

大模型可低成本替代人类实验,实现近最优统计推断。

Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data

  • 将大模型视为基于条件均值的统计估计器,理论证明其风险逼近贝叶斯最优。
  • 在满足条件时,大模型误差收敛至不可约方差加表示偏差,偏差受皮尔斯克不等式约束。
  • 提供校准方法与决策规则,适合需降低实验成本的研究者使用。

社会与行为科学中的量化研究依赖于昂贵、缓慢且易受抽样偏差影响的人类实验。本文表明,预训练大语言模型在平方损失下能诱导出条件期望的等风险估计器,实现受限功能风险等价:任何仅依赖数据条件均值的推断,其风险均与贝叶斯最优风险一致。我们将大模型形式化为基于i.i.d.数据的误设函数估计器 $T( ilde{P}_n)$,将估计误差分解为表示偏差 $ε_{\mathrm{rep}}$ 与优化误差,并证明在温和正则条件下,期望误差收敛至不可约总体方差加上平方表示偏差,其中表示偏差受皮尔斯克不等式限制。识别误差 $δ$ 会传播至有效偏差,抬高渐近风险下界。通过双向莱卡姆缺陷分析建立受限功能风险等价:前向缺陷渐近消失,反向缺陷精确为零。我们提供了有限样本浓度界与带明确决策规则的校准协议。结果是精确的可证明结论:经良好校准的大模型,在满足范围条件下,能达到条件均值相关推断的贝叶斯最优风险。实践中意味着,在条件满足且模型校准良好的情况下,大模型可用于诸多原本依赖人类实验的预测与决策任务,以更低成本实现近最优统计推断。

原文摘要 · Abstract (English)

Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias. Here we show that pretrained large language models induce risk-equivalent estimators of conditional expectations under squared loss, establishing restricted functional risk equivalence: under squared loss, the LLM induces an estimator whose risk matches the Bayes optimal risk for squared-loss prediction of conditional expectations for any inference that depends on the data only through the conditional mean. We formalize the LLM as a misspecified functional estimator $T(\hat{P}_n)$ trained on i.i.d.\ data, decompose the estimation error into representation bias $ε_{\mathrm{rep}}$ and optimization error, and prove that under mild regularity conditions the LLM's expected error converges to the irreducible population variance plus the squared representation bias, with the representation bias bounded by the Pinsker inequality. The identifiability error $δ$ propagates into the effective bias, inflating the asymptotic risk floor. We establish restricted functional risk equivalence via a bidirectional Le Cam deficiency analysis: the forward deficiency vanishes asymptotically while the reverse deficiency is exactly zero. We provide finite-sample concentration bounds and a calibration protocol with explicit decision rules. The result is a precise, provable statement: a well-calibrated LLM achieves the Bayes-optimal risk for conditional-mean-dependent inference, bounded by explicit scope conditions. In practical applications, this means that under satisfied conditions and well-calibrated models, large language models can be used in many prediction and decision-making tasks that originally relied on human experiments, approximating near-optimal statistical inference at lower cost.

大模型统计推断低成本实验条件均值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。