为病理图像虚拟染色设计专用相似性度量,提升模型评估与训练质量。
HAPS: Rethinking Image Similarity for Virtual Staining

- 提出基于病理图像特征的感知相似性度量HAPS,融合专家标注反馈。
- 在真实注册误差下测试,验证其对组织形态和生物标志物保留的敏感性。
- 用于清洗训练数据,显著提升虚拟染色模型性能,适合病理图像研究者。
虚拟染色(如H&E-IHC)是数字病理学中的新兴工具,可从常规切片合成目标染色图像,实现更快速、低成本的工作流程。然而,当前模型质量仍依赖SSIM、PSNR、LPIPS等通用指标,这些指标源于自然图像,与组织学数据的特异性不匹配,难以捕捉组织形态保真度与生物标志物表达模式。为此,本文将组织学图像相似性视为独立问题,系统评估多种全参考度量,并基于专家标注的H&E-IHC图像块对数据集进行分析。进一步考察度量在模拟真实注册误差的几何失真(平移、旋转、非刚性形变)下的敏感性。基于此,提出病理感知感知相似性(HAPS):使用预训练于病理数据的冻结编码器提取特征,在特征空间计算距离,并通过线性头聚合差异得到最终评分,与专家判断高度一致。最后,利用HAPS对MIST数据集的训练样本进行相似性量化并过滤低分样本,构建更清洁的训练集,经此优化的模型在虚拟染色任务中表现优于原始未筛选数据训练的模型。
原文摘要 · Abstract (English)
Virtual staining of histopathology images (e.g., H&E-IHC) is an emerging tool in digital pathology, enabling faster and cheaper workflows by synthesizing target stains from routinely acquired slides. Yet, the quality of virtual staining models is still predominantly assessed with generic metrics such as SSIM, PSNR, and LPIPS. Originally developed for natural images, these metrics are inherently misaligned with the domain-specific characteristics of histological data, failing to capture tissue morphology preservation and biomarker expression patterns. Consequently, a robust, domain-specific standard for quantifying similarity across diverse histological modalities remains a critical gap in the field. In this work, we formalize histology image similarity as a standalone problem and systematically evaluate a broad set of full-reference metrics against a dataset of H&E-IHC patch pairs annotated with expert similarity scores. We further analyze metrics sensitivity to controlled geometric distortions (shifts, rotations and non-rigid deformations) that mimic realistic registration errors between serial sections. Guided by these observations, we propose the Histology-Aware Perceptual Similarity (HAPS) metric. HAPS computes distances in the feature space of a frozen encoder pretrained on histopathology data, adding a linear head to aggregate feature-level differences into a final score that aligns with expert assessments. Finally, we demonstrate the practical value of HAPS for quality control of training data. By quantifying the similarity of training pairs in the MIST dataset and filtering low-scoring samples, we create a cleaner training set. Virtual staining models trained on this refined data outperform those trained on the original, unfiltered dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。