arXiv:2502.19231stat.MEcs.AI2025-02被引 6

用AI生成结果构建先验,实现可信的贝叶斯推断。

AI-Powered Bayesian Inference

  • 以AI生成数据为基线,构建狄利克雷过程先验。
  • 通过伪数据增强和优化,快速生成后验样本。
  • 适合需要量化不确定性的高风险决策场景。

生成式人工智能(GAI)虽不可完全信赖,但其输出的不确定性可被利用。本文提出在非参数贝叶斯框架下,将GAI生成结果作为数据生成分布的基线,建立狄利克雷过程先验。通过在样本外调整先验超参数,评估AI先验的信息量。后验推断通过在真实数据与由AI补全标签的伪数据组成的扩展数据集上计算随机化函数实现,可并行化且仅需优化即可快速生成独立同分布的后验样本。该方法实现了基于AI预测的连贯概率推断与不确定性量化。

原文摘要 · Abstract (English)

The advent of Generative Artificial Intelligence (GAI) has heralded an inflection point that changed how society thinks about knowledge acquisition. While GAI cannot be fully trusted for decision-making, it may still provide valuable information that can be integrated into a decision pipeline. Rather than seeing the lack of certitude and inherent randomness of GAI as a problem, we view it as an opportunity. Indeed, variable answers to given prompts can be leveraged to construct a prior distribution which reflects assuredness of AI predictions. This prior distribution may be combined with tailored datasets for a fully Bayesian analysis with an AI-driven prior. In this paper, we explore such a possibility within a non-parametric Bayesian framework. The basic idea consists of assigning a Dirichlet process prior distribution on the data-generating distribution with AI generative model as its baseline. Hyper-parameters of the prior can be tuned out-of-sample to assess the informativeness of the AI prior. Posterior simulation is achieved by computing a suitably randomized functional on an augmented data that consists of observed (labeled) data as well as fake data whose labels have been imputed using AI. This strategy can be parallelized and rapidly produces iid samples from the posterior by optimization as opposed to sampling from conditionals. Our method enables (predictive) inference and uncertainty quantification leveraging AI predictions in a coherent probabilistic manner.

贝叶斯推断AI先验不确定性量化生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。