不同病理模型导致FID评分差异30倍,提出标准化方法提升可比性。
HistoFID- Calibrating Frechet-distance evaluation across pathology foundation models

- 用自编码器内嵌的特征距离替代原始FID,实现跨模型可比性
- 标准化后跨模型变异系数降低89%(同组)和58%(跨组)
- 揭示编码器选择影响生成模型评估结论,适合病理图像质量评测
Frechet Inception Distance(FID)通过固定网络提取特征并拟合高斯分布来比较两组图像。在数字病理领域,通常用组织学基础模型替代Inception网络,假设域编码器能提供更合理的分数。我们发现该选择显著改变了结果:对于同一对切片集,六种常见编码器的原始弗雷歇距离相差约30倍,且排序不随嵌入维度变化,因此原始得分必须注明编码器。基于内部数据集(约50万张H&E与免疫组化切片,来自2,119张幻灯片)和公开TCGA BRCA数据集(100张幻灯片),我们基准测试了Inception-v3、Phikon-v2、CONCH、UNI2-h、Virchow2和Prov-GigaPath,在同数据集基线、跨数据集漂移、受控扰动、压缩、染色归一化及两个生成模型上的表现。将每项距离表示为相对于该编码器自身同组最低值的比率,恢复了可比性,使同组间变异系数下降约89%,跨组下降58%。编码器分为敏感组(CONCH、Phikon-v2、Inception-v3)和不变组(UNI2-h、Virchow2、Prov-GigaPath),这一划分决定了哪个生成模型被认为更真实,因此编码器可改变生成评估结论。在幻灯片级别,注意力池化编码器捕捉到拼贴距离无法察觉的局部组成信息,使距离上升约320倍。使用相同协议评估了专有病理编解码器TuroCompress,其在所有测试编解码器中以最小文件大小实现最高重建保真度。我们发布了标准化协议、各编码器扰动面板及特征提取结果。
原文摘要 · Abstract (English)
The Frechet Inception Distance (FID) compares two image sets by fitting a Gaussian to the features of a fixed network and measuring the distance between the two Gaussians. In digital pathology the Inception network is routinely replaced by a histology foundation model, on the assumption that a domain encoder gives a more meaningful score. We show that this choice changes the result. For one fixed pair of tile sets, the raw Frechet distance varies about thirty-fold across six common encoders, and the ordering does not follow embedding dimension, so a raw score cannot be read without naming the encoder. Using a held-out in-house cohort (about 500,000 H&E and immunohistochemistry tiles from 2,119 slides) and a public TCGA BRCA cohort (100 slides), we benchmark Inception-v3, Phikon-v2, CONCH, UNI2-h, Virchow2 and Prov-GigaPath across within-cohort baselines, cross-cohort drift, controlled perturbations, compression, stain normalization, and two generative models. Expressing each distance as a ratio to the encoder's own within-cohort floor restores comparability, cutting the across-encoder coefficient of variation by about 89% within cohort and 58% across cohorts. The encoders separate into a sensitive group (CONCH, Phikon-v2, Inception-v3) and an invariant group (UNI2-h, Virchow2, Prov-GigaPath), and this split decides which generative model is judged more realistic, so the encoder can change the conclusion of a generative evaluation. At the slide level, an attention-pooling encoder registers per-slide composition that a pooled patch distance cannot see, raising the distance about 320-fold on matched cohorts. Using the same protocol we evaluate TuroCompress, a proprietary pathology codec, which reaches the highest reconstruction fidelity at the smallest file size among codecs tested. We release the normalization protocol, the per-encoder perturbation panel, and the feature extracts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。