arXiv:2503.07962hep-phcs.LG2025-03被引 3

比较生成与判别方法在模拟推断中的表现,发现比率法更准更稳。

Discriminative versus Generative Approaches to Simulation-based Inference

  • 用判别分类和生成建模两种神经网络方法做参数推断
  • 在高斯和希格斯数据集上均有效提取参数,比率法精度更高
  • 两种方法都需集成或优化以减少训练波动,适合机器学习新手参考

粒子与核物理中的基础、涌现及现象学参数通常通过参数模板拟合确定。模拟用于填充直方图并与数据匹配,但该方法本质上是信息损失的,因直方图为分箱且低维。深度学习使无分箱、高维参数估计成为可能,通过神经似然(比)估计实现。本文对比了两种神经模拟推断(NSBI)方法:基于判别学习(分类)与基于生成建模的方法。两者在相同数据集上直接评估,超参数优化水平相当。除高斯数据集外,还使用了FAIR宇宙挑战中的希格斯玻色子数据集。结果表明,直接似然与似然比估计均可有效提取参数并获得合理不确定性。在研究的数值案例与超参数范围内,似然比方法更具准确性和/或精确性。两种方法均存在显著的网络训练差异,实际应用中需集成或其他缓解策略。

原文摘要 · Abstract (English)

Most of the fundamental, emergent, and phenomenological parameters of particle and nuclear physics are determined through parametric template fits. Simulations are used to populate histograms which are then matched to data. This approach is inherently lossy, since histograms are binned and low-dimensional. Deep learning has enabled unbinned and high-dimensional parameter estimation through neural likelihiood(-ratio) estimation. We compare two approaches for neural simulation-based inference (NSBI): one based on discriminative learning (classification) and one based on generative modeling. These two approaches are directly evaluated on the same datasets, with a similar level of hyperparameter optimization in both cases. In addition to a Gaussian dataset, we study NSBI using a Higgs boson dataset from the FAIR Universe Challenge. We find that both the direct likelihood and likelihood ratio estimation are able to effectively extract parameters with reasonable uncertainties. For the numerical examples and within the set of hyperparameters studied, we found that the likelihood ratio method is more accurate and/or precise. Both methods have a significant spread from the network training and would require ensembling or other mitigation strategies in practice.

模拟推断神经似然参数估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。