现有生成数据度量方法全有缺陷,影响真实应用。
Position: All Current Generative Fidelity and Diversity Metrics are Flawed
- 提出度量指标设计应满足的期望条件与一系列可靠性检验
- 实验证明所有现有度量在检测生成模型失败模式时均不靠谱
- 呼吁研究者优先改进度量而非模型,为实践提供使用指南
任何方法的发展和实际应用都受限于其可靠性度量能力。生成建模的流行凸显了优质合成数据度量的重要性。然而,已有研究发现当前度量存在诸多问题,例如对异常值不鲁棒、上下界模糊不清。本文提出合成数据度量应满足的一系列理想条件,并设计了一套精心选择的简单实验,旨在检测特定已知的生成建模失败模式。基于这些理想条件及检验结果,我们得出结论:所有现有的生成保真度与多样性度量均存在缺陷。这严重阻碍了合成数据的实际应用。我们的目标是说服研究社区将更多精力投入到度量开发上,而非模型本身。此外,通过分析现有度量的失效机制,我们为从业者提供了这些度量应如何(不应如何)使用的指导。
原文摘要 · Abstract (English)
Any method's development and practical application is limited by our ability to measure its reliability. The popularity of generative modeling emphasizes the importance of good synthetic data metrics. Unfortunately, previous works have found many failure cases in current metrics, for example lack of outlier robustness and unclear lower and upper bounds. We propose a list of desiderata for synthetic data metrics, and a suite of sanity checks: carefully chosen simple experiments that aim to detect specific and known generative modeling failure modes. Based on these desiderata and the results of our checks, we arrive at our position: all current generative fidelity and diversity metrics are flawed. This significantly hinders practical use of synthetic data. Our aim is to convince the research community to spend more effort in developing metrics, instead of models. Additionally, through analyzing how current metrics fail, we provide practitioners with guidelines on how these metrics should (not) be used.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。