发现生成模型评估中普遍存在评分虚高问题,提出新评分方法更贴近人类判断。
Grade Inflation in Generative Models
- 现有评分方法对合成数据质量评价偏高,存在系统性偏差。
- 新提出的Eden评分避免评分虚高,与人工判断更一致。
- 适合关注生成模型真实性能评估的研究者使用。
生成模型潜力巨大,但前提是其生成数据的评估可信。我们发现,许多用于比较二维分布的常用质量评分(如相关性、雅可比、地球移动者、相对熵)给出的结果优于实际水平,这一现象称为“评分虚高”。我们指出,所有对每个数据点同等重视的评分(称作“等点”评分)均存在此问题。为此提出“等密度”评分概念,并引入首个实例——Eden评分。实验表明,Eden评分避免了评分虚高,且与人类对拟合优度的感知更为一致。我们进一步揭示等密度评分与负阶瑞尼熵之间的联系,认为合理的等密度评分在生成模型及低维分布比较中可能优于等点评分。
原文摘要 · Abstract (English)
Generative models hold great potential, but only if one can trust the evaluation of the data they generate. We show that many commonly used quality scores for comparing two-dimensional distributions of synthetic vs. ground-truth data give better results than they should, a phenomenon we call the "grade inflation problem." We show that the correlation score, Jaccard score, earth-mover's score, and Kullback-Leibler (relative-entropy) score all suffer grade inflation. We propose that any score that values all datapoints equally, as these do, will also exhibit grade inflation; we refer to such scores as "equipoint" scores. We introduce the concept of "equidensity" scores, and present the Eden score, to our knowledge the first example of such a score. We found that Eden avoids grade inflation and agrees better with human perception of goodness-of-fit than the equipoint scores above. We propose that any reasonable equidensity score will avoid grade inflation. We identify a connection between equidensity scores and Rényi entropy of negative order. We conclude that equidensity scores are likely to outperform equipoint scores for generative models, and for comparing low-dimensional distributions more generally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。