用KL散度量化生成模型与真实分布的距离,给出可信的统计结论。
Statistical Inference for Generative Model Comparison
- 基于KL散度比较生成模型与真实分布的距离,无需调参。
- 在模拟数据上覆盖率达95%以上,比核方法检验力更强。
- 适用于图像、文本生成模型,可给出带置信度的对比结果。
生成模型在诸多应用中取得显著成功,但其评估仍缺乏严谨的不确定性量化。本文提出一种方法,用于比较不同生成模型与测试样本真实分布的接近程度。特别地,采用Kullback-Leibler(KL)散度度量生成模型与未知测试分布之间的距离,因为KL散度无需像基于RKHS距离那样依赖核函数调参,且是唯一能实现关键抵消以支持不确定性量化f-散度。此外,我们将方法扩展至条件生成模型,并利用埃奇沃思展开处理小样本情形。在具有已知真值的模拟数据上,我们的方法实现了有效的覆盖率,且相比核基方法具有更高的检验功效。应用于图像和文本数据集上的生成模型时,该程序得出的结果与基准指标一致,但附带统计置信度。
原文摘要 · Abstract (English)
Generative models have achieved remarkable success across a range of applications, yet their evaluation still lacks principled uncertainty quantification. In this paper, we develop a method for comparing how close different generative models are to the underlying distribution of test samples. Particularly, our approach employs the Kullback-Leibler (KL) divergence to measure the distance between a generative model and the unknown test distribution, as KL requires no tuning parameters such as the kernels used by RKHS-based distances, and is the only $f$-divergence that admits a crucial cancellation to enable the uncertainty quantification. Furthermore, we extend our method to comparing conditional generative models and leverage Edgeworth expansions to address limited-data settings. On simulated datasets with known ground truth, we show that our approach realizes effective coverage rates, and has higher power compared to kernel-based methods. When applied to generative models on image and text datasets, our procedure yields conclusions consistent with benchmark metrics but with statistical confidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。