用回归投影加速模拟推断,高效且可并行计算。
Simulation-Based Inference via Regression Projection and Batched Discrepancies
- 通过一次拟合线性回归,用小批量模拟估算参数后验
- 伪后验随样本量增大一致收敛,批大小增大会集中于识别集
- 适合高维模拟器的快速参数估计,但受限于低信息汇总
我们分析了一种轻量级的基于模拟的推断方法,该方法仅通过观测数据的回归投影来推断模拟器参数。在一次拟合代理线性回归后,对提议的参数值进行小批量模拟,根据批次残差差异分配核权重,生成自归一化的伪后验。该方法简单、可并行,仅需访问已拟合的回归系数,无需原始观测。我们将构造形式化为对模拟随机性的总体目标的加权重要性采样近似,证明了随着参数抽样次数增加,方法具有一致性,并建立了从有限样本中估计代理回归的稳定性。随后,我们刻画了当批大小增加、带宽缩小下的渐近集中行为,表明伪后验会集中在由所选投影决定的可识别集合上,从而阐明了方法产生点识别还是集识别的条件。在可解析的非线性模型和使用DREAMS模拟套件的宇宙学校准任务上的实验,展示了回归投影的计算优势以及由低信息摘要引发的可识别性局限。
原文摘要 · Abstract (English)
We analyze a lightweight simulation-based inference method that infers simulator parameters using only a regression-based projection of the observed data. After fitting a surrogate linear regression once, the procedure simulates small batches at the proposed parameter values and assigns kernel weights based on the resulting batch-residual discrepancy, producing a self-normalized pseudo-posterior that is simple, parallelizable, and requires access only to the fitted regression coefficients rather than raw observations. We formalize the construction as an importance-sampling approximation to a population target that averages over simulator randomness, prove consistency as the number of parameter draws grows, and establish stability in estimating the surrogate regression from finite samples. We then characterize the asymptotic concentration as the batch size increases and the bandwidth shrinks, showing that the pseudo-posterior concentrates on an identified set determined by the chosen projection, thereby clarifying when the method yields point versus set identification. Experiments on a tractable nonlinear model and on a cosmological calibration task using the DREAMS simulation suite illustrate the computational advantages of regression-based projections and the identifiability limitations arising from low-information summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。