用像素级染色准确率评估虚拟免疫组化图像质量,更贴近病理诊断需求。
Building Trust in Virtual Immunohistochemistry: Automated Assessment of Image Quality
- 基于颜色解卷积生成虚拟与真实IHC的阳性像素掩码
- 配对模型如PyramidPix2Pix在染色准确性上最优,优于无配对模型
- 发现局部切片评估无法发现全切片性能下降,需全局评估
深度学习模型可从苏木精-伊红(H&E)图像生成虚拟免疫组化(IHC)染色,提供一种可扩展且低成本的替代方案。然而,现有基于纹理和分布的图像保真度指标仅衡量视觉相似性,而非染色准确性。本文提出一种自动化的、以准确性为基础的框架,用于评估十六种配对或非配对图像转换模型的图像质量。通过颜色解卷积,生成各虚拟IHC模型预测的棕褐色(IHC阳性)像素掩码,并使用分割掩码计算骰子系数(Dice)、交并比(IoU)和豪斯多夫距离等染色准确率指标,无需专家人工标注即可量化像素级标签正确性。结果表明,传统保真度指标(如FID、PSNR、SSIM)与染色准确率及病理科医生评估的相关性极低。配对模型(如PyramidPix2Pix和AdaptiveNCE)表现最佳,而基于扩散模型和GAN的非配对模型在准确标记IHC阳性像素方面可靠性较差。此外,全切片图像(WSI)分析揭示了局部切片评估中未察觉的性能下降,凸显了在全切片层面建立基准的重要性。该框架为虚拟IHC模型的质量评估提供了可复现的方法,是推动其走向临床常规应用的关键一步。
原文摘要 · Abstract (English)
Deep learning models can generate virtual immunohistochemistry (IHC) stains from hematoxylin and eosin (H&E) images, offering a scalable and low-cost alternative to laboratory IHC. However, reliable evaluation of image quality remains a challenge as current texture- and distribution-based metrics quantify image fidelity rather than the accuracy of IHC staining. Here, we introduce an automated and accuracy grounded framework to determine image quality across sixteen paired or unpaired image translation models. Using color deconvolution, we generate masks of pixels stained brown (i.e., IHC-positive) as predicted by each virtual IHC model. We use the segmented masks of real and virtual IHC to compute stain accuracy metrics (Dice, IoU, Hausdorff distance) that directly quantify correct pixel - level labeling without needing expert manual annotations. Our results demonstrate that conventional image fidelity metrics, including Frechet Inception Distance (FID), peak signal-to-noise ratio (PSNR), and structural similarity (SSIM), correlate poorly with stain accuracy and pathologist assessment. Paired models such as PyramidPix2Pix and AdaptiveNCE achieve the highest stain accuracy, whereas unpaired diffusion- and GAN-based models are less reliable in providing accurate IHC positive pixel labels. Moreover, whole-slide images (WSI) reveal performance declines that are invisible in patch-based evaluations, emphasizing the need for WSI-level benchmarks. Together, this framework defines a reproducible approach for assessing the quality of virtual IHC models, a critical step to accelerate translation towards routine use by pathologists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。