arXiv:2512.23184cs.AIecon.EM2025-12

用语言模型的信念分布提升数据效率,让生成结果更准更快。

From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research

  • 从模型输出概率中提取信念分布,比直接用选择结果更有效。
  • 在有限运行次数下,信念分布预测真实选择效果更好,计算量少20倍。
  • 适合做行为模拟、需求估计等需高效生成数据的研究者使用。

大型语言模型(LLMs)被广泛用于模拟人类行为,但通常将模型输出(即‘模型选择’)当作单一数据点,未能充分利用其概率特性。本文提出并形式化了‘模型信念’——一种基于模型生成时的逐标记概率构建的信念分布,用于捕捉单次生成中对不同选项的置信度。作者证明,模型信念在渐近意义上等价于模型选择的均值(非平凡性质),但具有更低方差和更快收敛速度,是更高效的统计估计器。对于下游应用中常见的信念或选择的光滑函数,类似性质同样成立。通过一个需求估计研究,实验表明:在有限生成次数下,模型信念比模型选择更能解释和预测真实选择,且达到足够精度所需的计算量减少约20倍。结果支持将模型信念作为从语言模型生成数据中提取信息的标准方法。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output ("model choice") as a single data point underutilizes the information inherent to the probabilistic nature of LLMs. This paper introduces and formalizes "model belief," a measure derived from an LLM's token-level probabilities that captures the model's belief distribution over choice alternatives in a single generation run. The authors prove that model belief is asymptotically equivalent to the mean of model choices (a non-trivial property) but forms a more statistically efficient estimator, with lower variance and a faster convergence rate. Analogous properties are shown to hold for smooth functions of model belief and model choice often used in downstream applications. The authors demonstrate the performance of model belief through a demand estimation study, where an LLM simulates consumer responses to different prices. In practical settings with limited numbers of runs, model belief explains and predicts ground-truth model choice better than model choice itself, and reduces the computation needed to reach sufficiently accurate estimates by roughly a factor of 20. The findings support using model belief as the default measure to extract more information from LLM-generated data.

大模型行为模拟数据效率信念分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。