arXiv:2512.01107cs.AIecon.EM2025-12

将大模型生成内容视为带主观偏见的先验,而非真实数据。

Foundation Priors

  • 把生成结果看作用户先验与模型分布的加权混合
  • 引入信任参数控制合成数据权重,避免误当真实数据
  • 适合做模型修正、实验设计等需要主观先验的场景

大语言模型生成的内容常被当作真实数据使用,但本文提出‘基础先验’概念:生成结果并非客观观测,而是反映模型学习模式与用户主观预期、偏见的联合产物。通过显式建模用户对数据分布的预期、提示工程过程及对模型的信任程度,本文将基础先验定义为用户原始先验的指数倾斜化贝叶斯更新,其中信任参数决定合成数据的影响权重。该框架可无缝融入标准统计与计量分析流程,应用于复杂模型优化、潜变量推断、实验设计以及随机系数和部分线性模型的增强。通过将生成内容视为结构化的主观先验,而非实证观察,该方法为在实证研究中合理使用大模型提供了理论依据,有效防止合成‘事实’与真实数据混淆。

原文摘要 · Abstract (English)

Foundation models, and in particular large language models, can generate highly informative responses, prompting growing interest in using these ''synthetic'' outputs as data in empirical research and decision-making. This paper introduces the idea of a foundation prior, which shows that model-generated outputs are not as real observations, but draws from the foundation prior induced prior predictive distribution. As such synthetic data reflects both the model's learned patterns and the user's subjective priors, expectations, and biases. We model the subjectivity of the generative process by making explicit the dependence of synthetic outputs on the user's anticipated data distribution, the prompt-engineering process, and the trust placed in the foundation model. We derive the foundation prior as an exponential-tilted, generalized Bayesian update of the user's primitive prior, where a trust parameter governs the weight assigned to synthetic data. We then show how synthetic data and the associated foundation prior can be incorporated into standard statistical and econometric workflows, and discuss their use in applications such as refining complex models, informing latent constructs, guiding experimental design, and augmenting random-coefficient and partially linear specifications. By treating generative outputs as structured, explicitly subjective priors rather than as empirical observations, the framework offers a principled way to harness foundation models in empirical work while avoiding the conflation of synthetic ''facts'' with real data.

大模型先验建模合成数据主观偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。