arXiv:2604.20723cs.LGcs.AI2026-04

用单点模拟训练模型,大幅提升层次化推断效率

Tokenised Flow Matching for Hierarchical Simulation Based Inference

论文配图:Tokenised Flow Matching for Hierarchical Simulation Based Inference
图 1 · 摘自论文原文
  • 通过似然分解,仅需单点模拟即可训练模型
  • 在传染病与流体动力学模型中实现校准后验且降低计算成本
  • 适合需要高效层次推断的科研与工程场景

模拟器评估成本是模拟基础推断(SBI)的关键瓶颈。在具有共享全局参数和可交换站点级参数与观测值的层次结构中,可利用此特性提升模拟效率。现有层次化SBI方法虽对后验进行因子分解,但仍需对每个训练样本模拟多个站点;我们改用似然因子分解(LF),从单站点模拟中训练。在LF采样中,学习每个站点的神经模拟器代理,再合成多站点观测以摊销完整层次后验的推断。基于此,我们提出用于后验估计的分块流匹配方法(TFMPE),支持函数型观测的似然因子分解。为实现系统评估,我们引入一个层次化SBI基准。在该基准及真实感染疾病与计算流体动力学模型上验证了TFMPE,结果表明其能获得校准良好的后验分布,同时显著降低计算开销。

原文摘要 · Abstract (English)

The cost of simulator evaluations is a key practical bottleneck for Simulation Based Inference (SBI). In hierarchical settings with shared global parameters and exchangeable site-level parameters and observations, this structure can be exploited to improve simulation efficiency. Existing hierarchical SBI approaches factorise the posterior yet still simulate across multiple sites per training sample; We instead explore likelihood factorisation (LF) to train from single-site simulations. In LF sampling we learn a per-site neural surrogate of the simulator and then assemble synthetic multi-site observations to amortise inference for the full hierarchical posterior. Building on this, we propose Tokenised Flow Matching for Posterior Estimation (TFMPE), a tokenised flow matching approach that supports function-valued observations through likelihood factorisation. To enable systematic evaluation, we introduce a benchmark for hierarchical SBI. We validate TFMPE on this benchmark and on realistic infectious disease and computational fluid dynamics models, finding well-calibrated posteriors while reducing computational cost.

层次推断模拟器流匹配高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。