arXiv:2601.21979cs.LG2026-01

用随机丢弃法评估FID在医学图像中的可信度,发现其结果波动与数据分布偏移相关。

Evaluating the trustworthiness of the Fréchet Inception Distance with stochastic embedding representations

  • 通过蒙特卡洛丢弃法计算FID和特征嵌入的预测方差。
  • FID的方差大小与输入数据偏离训练分布的程度呈一定相关性。
  • 为医学图像生成质量评估提供了可信度判断新思路,适合关注生成模型可靠性的研究者。

从预训练模型中获取的特征嵌入广泛应用于深度学习的医疗领域,用于评估数据集特性,例如合成医学图像的质量。弗雷谢特初始距离(FID)是一种流行的合成图像质量指标,依赖于InceptionV3模型在ImageNet1K(自然图像)上预训练的假设,即数据特征可被有效检测和编码。尽管众所周知该方法在医学图像应用中效果较差,但其无法捕捉图像特征差异的程度尚不明确。本文采用蒙特卡洛丢弃法,计算FID及特征嵌入模型潜在表示的预测方差。结果显示,预测方差的大小与测试输入(经不同强度增强的ImageNet1K验证集及其他外部数据集)相对于训练数据的分布外程度存在不同程度的相关性,为使用这些方差作为FID可信度指标提供了依据。

原文摘要 · Abstract (English)

Feature embeddings acquired from pretrained models are widely used in medical applications of deep learning to assess the characteristics of datasets; e.g. to determine the quality of synthetic, generated medical images. The Fréchet Inception Distance (FID) is one popular synthetic image quality metric that relies on the assumption that the characteristic features of the data can be detected and encoded by an InceptionV3 model pretrained on ImageNet1K (natural images). While it is widely known that this makes it less effective for applications involving medical images, the extent to which the metric fails to capture meaningful differences in image characteristics is not obviously known. Here, we use Monte Carlo dropout to compute the predictive variance in the FID as well as a supplemental estimate of the predictive variance in the feature embedding model's latent representations. We show that the magnitudes of the predictive variances considered exhibit varying degrees of correlation with the extent to which test inputs (ImageNet1K validation set augmented at various strengths, and other external datasets) are out-of-distribution relative to its training data, providing some insight into the effectiveness of their use as indicators of the trustworthiness of the FID.

FID可信度评估医学图像不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。