arXiv:2605.29335cs.CVcs.AI2026-05中稿 · ICML

FID评估结果受真实数据分布几何形状影响,不能单独判断生成质量。

Rethinking FID Through the Geometry of the Reference Dataset

  • 分析真实数据集的分布密度与有效秩对FID的影响
  • 密集数据集使FID更可信,分散数据集可能导致误判
  • 建议结合数据几何特性使用FID,适合模型评估者参考

Fréchet Inception Distance(FID)被广泛用于图像生成器的评估,但更低的FID并不总意味着更好的样本质量。我们发现这种不一致部分源于参考数据集的几何结构。在六个数据集上的控制实验表明,分布密度和有效秩显著影响FID随生成质量提升的变化趋势。集中分布的数据集通常产生更有利的FID变化,而更分散的数据集可能在生成质量提高时导致FID反而恶化。通过精度与召回的归因分析,以及在不同特征空间和距离度量下的消融实验,均支持这一结论。结果表明,分布度量应结合参考数据集的几何特性进行解释,以实现更可靠的基准测试。

原文摘要 · Abstract (English)

Fréchet Inception Distance (FID) is widely used to evaluate image generators, yet lower FID does not always correspond to better sample quality. We show that this mismatch depends in part on the geometry of the reference dataset. In a controlled study across six datasets, distributional density and effective rank significantly explain how FID changes as sample quality improves. Concentrated datasets tend to yield more favorable FID trends, whereas more dispersed datasets can make FID worsen despite better samples. Attribution to precision and recall and ablations with alternative feature spaces and distances support the same conclusion. These results suggest that distributional metrics should be interpreted together with the geometry of the reference dataset for more reliable benchmarking.

FID生成评估数据分布图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。