arXiv:2502.17160cs.CVcs.LG2025-02被引 15

FID评估生成图像在眼底影像中可能失准,需以实际任务表现为准。

A Pragmatic Note on Evaluating Generative Models with Fréchet Inception Distance for Retinal Image Synthesis

  • 用真实任务如分类分割来评估生成数据价值
  • 发现FID在眼底图像上与实际效果不一致
  • 适合医学图像生成的开发者和临床研究者

Fréchet Inception Distance(FID)使用预训练的ImageNet Inception-v3网络计算特征向量的2-Wasserstein距离,常被用作生成模型的先进评估指标。该方法假设Inception-v3特征服从多元高斯分布。尽管在多数图像合成任务中能有效衡量生成数据与真实数据的接近程度,但在生物医学生成模型中,主要目标通常是生成带标注的数据以扩充训练集。因此,评估生成模型的金标准应是将其合成数据用于下游任务(如分类、分割)并观察性能提升。本文以眼底成像(包括彩色眼底照片和光学相干断层扫描)为例,揭示了FID及其变体在分类和分割任务中与实际评估目标存在偏差。我们指出,此类指标在更广泛的生物医学成像及下游任务中可能存在潜在问题。

原文摘要 · Abstract (English)

Fréchet Inception Distance (FID), computed with an ImageNet pretrained Inception-v3 network, is widely used as a state-of-the-art evaluation metric for generative models. It assumes that feature vectors from Inception-v3 follow a multivariate Gaussian distribution and calculates the 2-Wasserstein distance based on their means and covariances. While FID effectively measures how closely synthetic data match real data in many image synthesis tasks, the primary goal in biomedical generative models is often to enrich training datasets ideally with corresponding annotations. For this purpose, the gold standard for evaluating generative models is to incorporate synthetic data into downstream task training, such as classification and segmentation, to pragmatically assess its performance. In this paper, we examine cases from retinal imaging modalities, including color fundus photography and optical coherence tomography, where FID and its related metrics misalign with task-specific evaluation goals in classification and segmentation. We highlight the limitations of using various metrics, represented by FID and its variants, as evaluation criteria for these applications and address their potential caveats in broader biomedical imaging modalities and downstream tasks.

图像生成医学影像评估指标眼底成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。