arXiv:2608.03990cs.LG2026-08

用病理专用指标评估生成病理图像质量,发现数据多样性比画质更重要。

Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation

  • 改用病理预训练模型计算FID和IS,提升评估准确性。
  • 新指标与细胞分割性能相关性达0.61,原指标仅0.07。
  • 生成数据多样性对模型性能影响更大,优于单图画质提升。

合成病理图像生成有望缓解计算病理学中的数据稀缺问题,但现有评估方法可能无法充分衡量其医学应用质量。本文研究并改进了现有评估指标的局限性,提出基于领域特定指标和下游任务验证的评估方法。发现传统评估指标如弗雷谢特激活距离(FID)和判别分数(IS)在病理图像上存在不足,因其依赖ImageNet预训练特征提取器。为此,我们建议使用在数字病理数据集上预训练的基础模型改进FID和IS,并引入基于精确率-召回率的指标作为补充。基于四个基准数据集,采用两阶段训练的条件去噪扩散模型,生成具有系统性质量差异的合成数据集。同时,通过聚合交并比(AJI+)和骰子系数等指标,测量合成数据质量与下游核分割性能的相关性。结果表明,病理专用指标具备更强区分能力:改进后的判别分数与下游任务性能相关性为r=0.6096(p=0.0122),显著高于原始IS的r=0.0708(p=0.7944)。观察显示,增加生成数据的多样性比提升单张图像视觉保真度对分割模型性能有更高正向影响。

原文摘要 · Abstract (English)

Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data quality for medical applications. This work investigates and addresses limitations in existing evaluation metrics, investigating an approach for assessing synthetic histopathology image quality through domain-specific metrics and downstream task validation. We show that conventional synthetic data evaluation metrics such as Frechet Inception Distance (FID) and Inception Score (IS) may have limitations when applied to histopathology images due to their reliance on ImageNet-pretrained feature extractors. To address these limitations, we propose for consideration modified FID and IS approaches utilizing foundation models pretrained on digital pathology datasets, supplemented by precision-recall based metrics as part of an additional quality assessment. Using conditional denoising diffusion models trained on four benchmark datasets, with a two-step training approach, we generated synthetic datasets with systematically varied quality characteristics. We also measured the correlation between the synthetic data quality metrics with downstream nuclei segmentation performance using common metrics including the aggregated Jaccard index (AJI+) and the Dice coefficient. The study results suggest that pathology-specific metrics may provide improved discriminative power. Specifically, the modified Inception Score indicates higher correlation with downstream task performance (r=0.6096 with AJI+, p=0.0122), compared to the original IS (r=0.0708, p=0.7944). Our observations indicate that increasing the variety of generated training data has a higher positive correlation with segmentation model performance than improving the visual fidelity of individual generated images.

病理生成扩散模型图像评估数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。