FID评估有盲区,新方法ZID能精准检测生成模型的分布偏差。
What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation

- 提出ZID诊断框架,融合位置与离散度敏感指标,突破传统单标量评估局限。
- 在ImageNet上,仅匹配均值协方差的图像FID达24.7,远优于真实图像的58.6。
- 可识别模式崩溃导致的过低离散度,适合模型质量诊断与调优参考。
生成模型常以弗雷谢特初始距离(FID)和核初始距离(KID)排序,但其仅基于前两阶矩的摘要特性可能遗漏分布差异,且单一标量差距无法校准采样变异。FID的矩约束带来实际后果:在ImageNet上,仅优化匹配参考Inception均值与协方差的不可识别图像,其FID为24.7,而真实图像的FID为58.6(越低越好)。此外,FID与KID为标量差异,交换样本后不变,无法反映离散度变化方向——如模式崩溃导致的欠离散或过离散。本文提出ZID(Z-解析集成诊断),结合秩图(RISE)的六种标准化位置与离散度敏感分支及两种带宽的高斯核(GPK)。ZID不依赖单一标量,而是输出三个关联结果:用于排序偏离程度的指数、检验分布相等性的置换p值,以及用于诊断的符号化离散度读数。在受控实验中,ZID可检测广泛偏离,其得分随严重程度递增,包括FID保持平坦甚至反转的情形。在DiT-XL/2与SiT-XL/2引导强度扫描中,ZID成功检测出与真实数据的偏离,其符号读数将高引导下的多样性崩溃明确标记为欠离散。
原文摘要 · Abstract (English)
Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable images optimized only to match the reference Inception mean and covariance obtain FID $24.7$ versus $58.6$ for held-out real images (lower is better). Moreover, FID and KID are scalar discrepancies that are unchanged when the two samples are exchanged and therefore do not encode the direction of a dispersion change: under-dispersion, as can occur in mode collapse, versus over-dispersion. We introduce \textbf{ZID} (\emph{Z-resolved Integrated Diagnostic}), which combines six standardized location- and dispersion-sensitive arms from a rank graph (RISE) and Gaussian kernels (GPK at two bandwidths). Rather than asking one scalar to serve incompatible roles, ZID reports three linked outputs: an index for ranking departure magnitude, a permutation $p$-value for testing distributional equality, and a signed dispersion readout for diagnosis. In controlled experiments, ZID detects a broad range of departures, and its score tracks increasing severity along the corresponding sweeps, including cases in which FID is flat or reversed. On DiT-XL/2 and SiT-XL/2 guidance sweeps, ZID detects departure from real data, and its signed readout labels the high-guidance diversity collapse as under-dispersion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。