提出两种新评估指标,让生成模型质量判断更可靠、可解释。
Enhanced Generative Model Evaluation with Clipped Density and Coverage
- 通过截断样本贡献和最近邻半径,抑制异常样本干扰。
- 得分随劣质样本比例线性下降,可直接理解为优质样本占比。
- 在真实与合成数据上均优于现有方法,适合高风险场景评估。
尽管生成模型近年进展显著,但在关键应用中仍受限于难以可靠评估生成样本的质量。质量包含保真度与覆盖率两个互补维度。现有评价指标因缺乏校准或对异常值不鲁棒,常给出不可靠、难解释的数值。为此,本文提出两种新指标:截断密度(Clipped Density)与截断覆盖率(Clipped Coverage)。通过截断单个样本贡献及最近邻球半径来衡量保真度,有效防止分布外样本扭曲整体评估结果。经理论与实证校准,指标得分随劣质样本比例呈线性下降,可直观解释为优质样本的比例。在合成与真实数据集上的大量实验表明,该方法在鲁棒性、敏感性和可解释性方面均优于现有方法。
原文摘要 · Abstract (English)
Although generative models have made remarkable progress in recent years, their use in critical applications has been hindered by an inability to reliably evaluate the quality of their generated samples. Quality refers to at least two complementary concepts: fidelity and coverage. Current quality metrics often lack reliable, interpretable values due to an absence of calibration or insufficient robustness to outliers. To address these shortcomings, we introduce two novel metrics: Clipped Density and Clipped Coverage. By clipping individual sample contributions, as well as the radii of nearest neighbor balls for fidelity, our metrics prevent out-of-distribution samples from biasing the aggregated values. Through analytical and empirical calibration, these metrics demonstrate linear score degradation as the proportion of bad samples increases. Thus, they can be straightforwardly interpreted as equivalent proportions of good samples. Extensive experiments on synthetic and real-world datasets demonstrate that Clipped Density and Clipped Coverage outperform existing methods in terms of robustness, sensitivity, and interpretability when evaluating generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。