用一维统计方法高效评估高维生成模型,速度快且敏感度不输传统方法。
Refereeing the Referees: Evaluating Two-Sample Tests for Validating Generators in Precision Sciences
- 基于一维积分概率测度设计可并行计算的检验方法
- 在5~100维高维数据上表现接近多变量方法,但计算成本大幅降低
- 适合粒子物理等高维科学场景中的生成模型验证,可作基准工具
我们提出一种稳健的方法,评估非参数两样本检验在高维生成模型(如粒子物理)中的性能与计算效率。研究聚焦于基于一维积分概率测度的检验:切片 Wasserstein 距离、柯尔莫哥洛夫-斯米尔诺夫统计量均值,以及新提出的切片柯尔莫哥洛夫-斯米尔诺夫统计量。这些指标可并行计算,能快速可靠地估计零假设下的分布。我们还对比了最近提出的无偏 Fréchet 高斯距离和无偏二次最大均值差异(使用四次多项式核)。实验涵盖5、20、100维的协方差高斯与高斯混合分布,以及来自 JetNet 数据集的胶子喷注粒子级与喷注级特征。结果表明,一维方法在敏感度上可媲美多变量指标,但计算开销显著更低,适用于高维科学生成模型的评估。该方法为模型比较提供高效标准化工具,亦可作为更先进测试(包括基于机器学习的方法)的基准。
原文摘要 · Abstract (English)
We propose a robust methodology to evaluate the performance and computational efficiency of non-parametric two-sample tests, specifically designed for high-dimensional generative models in scientific applications such as in particle physics. The study focuses on tests built from univariate integral probability measures: the sliced Wasserstein distance and the mean of the Kolmogorov-Smirnov statistics, already discussed in the literature, and the novel sliced Kolmogorov-Smirnov statistic. These metrics can be evaluated in parallel, allowing for fast and reliable estimates of their distribution under the null hypothesis. We also compare these metrics with the recently proposed unbiased Fréchet Gaussian Distance and the unbiased quadratic Maximum Mean Discrepancy, computed with a quartic polynomial kernel. We evaluate the proposed tests on various distributions, focusing on their sensitivity to deformations parameterized by a single parameter $ε$. Our experiments include correlated Gaussians and mixtures of Gaussians in 5, 20, and 100 dimensions, and a particle physics dataset of gluon jets from the JetNet dataset, considering both jet- and particle-level features. Our results demonstrate that one-dimensional-based tests provide a level of sensitivity comparable to other multivariate metrics, but with significantly lower computational cost, making them ideal for evaluating generative models in high-dimensional settings. This methodology offers an efficient, standardized tool for model comparison and can serve as a benchmark for more advanced tests, including machine-learning-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。