用迭代后验平均法更准估计模拟器参数分布。
Source Distribution Estimation by Posterior Averaging
- 通过交替更新后验与源分布,动态优化参数估计。
- 在洛特卡-沃尔泰拉模型上,新方法误差低至0.64-0.68(原方法均>0.96)。
- 适合初始先验宽或错配的复杂模拟场景,提升鲁棒性。
基于模拟的科学常需估计模拟器参数的分布,使其输出能匹配真实观测数据,即源分布估计(SDE)问题。现有方法依赖一次训练的似然代理,目标仅基于代理而非真实模拟器,导致参数空间中未覆盖区域性能下降。本文提出基于期望最大化的方法:E步利用当前源估计的最新模拟训练可扩展后验,M步将源重新拟合为该后验在观测数据上的平均。给出两种参数化:(1) 分离的源与后验流;(2) 共享条件流。在三个基准任务上测试,涵盖广义与错误设定的初始先验。两种方法均优于固定代理及迭代变体,尤其在洛特卡-沃尔泰拉模型中,所有基线方法数据空间C2ST均高于0.96,而本方法在四种初始先验设置中有三种达到0.64–0.68。
原文摘要 · Abstract (English)
Simulation-based science often requires a distribution over simulator parameters whose push-forward reproduces a set of real observations: this is the source distribution estimation (SDE) problem. Existing methods fit the source against a likelihood surrogate trained once from a fixed proposal prior. Their objective is therefore stated only in terms of the surrogate instead of the true simulator, which may fail for inaccurate areas in parameter space where the surrogate was never trained. We instead solve SDE by expectation maximization: an E-step trains an amortized posterior on fresh simulations from the current source estimate, and an M-step refits the source to the average of that posterior over the observed data. We give two parameterizations, (1) separate source and posterior flows and (2) a single shared conditional flow. We evaluate our method on three benchmark tasks under both broad and misspecified initial priors. Both improve on existing fixed surrogate approaches and on iterated variants of each, most clearly on Lotka--Volterra, where no baseline falls below 0.96 data-space C2ST while our methods reach 0.64-0.68 in three of four initial-prior settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。