用新方法让生成模型推断更可信,保证结果真实有效
Trustworthy scientific inference with generative models
- 将生成模型的后验分布转化为可验证的置信区间
- 在数据分布变化时仍能准确识别隐藏参数
- 适合需要可靠推断的物理与观测科学领域
生成式人工智能擅长从训练样本中学习模式并生成复杂数据结构(如文本、图像、视频)。在多个科学领域,研究者正将生成模型应用于“反问题”,直接从观测数据预测隐藏参数,并提供不确定性度量。尽管这类基于后验的方法能处理难以计算似然的情况和大规模研究,但即使模型无误设,仍可能产生有偏或过度自信的结论。本文提出频率-贝叶斯协议(Frequentist-Bayes, FreB),一种数学严谨的解决方案,可将生成模型产生的后验概率分布重塑为局部有效的置信区域,在期望概率下始终包含真实参数,且当训练与目标数据对齐时达到最小尺寸。我们通过物理科学中的多个案例验证了FreB的有效性:在数据分布偏移下识别未知源、调和竞争理论模型之间的矛盾,以及缓解观测研究中的选择偏差与系统误差。通过提供可验证的可靠性保障和可解释诊断工具,FreB使在无法直接评估似然或计算成本过高的领域实现可信科学推断成为可能。
原文摘要 · Abstract (English)
Generative artificial intelligence (AI) excels at producing complex data structures (text, images, videos) by learning patterns from training examples. Across scientific disciplines, researchers are now applying generative models to "inverse problems" to directly predict hidden parameters from observed data along with measures of uncertainty. While these predictive or posterior-based methods can handle intractable likelihoods and large-scale studies, they can also produce biased or overconfident conclusions even without model misspecifications. We present a solution with Frequentist-Bayes (FreB), a mathematically rigorous protocol that reshapes AI-generated posterior probability distributions into (locally valid) confidence regions that consistently include true parameters with the expected probability, while achieving minimum size when training and target data align. We demonstrate FreB's effectiveness by tackling diverse case studies in the physical sciences: identifying unknown sources under dataset shift, reconciling competing theoretical models, and mitigating selection bias and systematics in observational studies. By providing validity guarantees with interpretable diagnostics, FreB enables trustworthy scientific inference across fields where direct likelihood evaluation remains impossible or prohibitively expensive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。