对比6种生成模型,评估无配对虚拟染色的可靠性与不确定性。
Towards Reliable AI-Based Histological Staining: A Systematic Study of Scaling and Uncertainty in Unpaired Generative Models

- 构建新数据集,系统测试6类无配对模型在54种配置下的表现
- 发现感知质量、任务误差与模型一致性三者独立,需联合评估
- 首次量化无监督模型的预测不确定性,适合病理AI开发者参考
肝纤维化是慢性肝病长期预后的关键指标,通过胶原含量的组织学估计进行分期。伊红苏木精(H&E)是常规染色,而Sirius Red(SR)虽提供标准定量读数(胶原比例面积,CPA),但并非所有临床中心都采用,且耗时耗材。基于AI的虚拟染色可从H&E生成SR,但无监督模型的系统性基准和预测不确定性尚未量化,尽管视觉上逼真的输出未必忠实反映组织结构。因此,我们在新发布的配对小鼠肝组织H&E到SR数据集上,系统评估了六种无监督图像到图像架构(基于GAN和扩散模型),共54种缩放配置。每种配置在感知、分布和任务特定三个维度及盲法专家阅读研究中综合评估;最优家族模型被重训练为深度集成模型,首次系统比较无监督染色到染色架构的主观不确定性。跨模型家族间,感知质量、任务误差与集成一致性为独立评价维度:GAN类模型在感知指标上聚类紧密,但在任务误差和集成一致性上差异显著;而扩散模型(CycleDiffusion)在三方面均表现迥异。单一指标无法捕捉这些差异,可靠虚拟染色需联合报告与选择三类指标。数据集、分块流程、模型与评估代码均已公开。
原文摘要 · Abstract (English)
Liver fibrosis, the principal predictor of long-term outcome in chronic liver disease, is staged from histological estimates of collagen content. Sirius Red (SR) provides the standard quantitative readout (collagen proportionate area, CPA) but is not acquired at every clinical centre and consumes tissue, time, and reagent cost beyond the routine Hematoxylin and eosin (H&E) stain. AI-based virtual staining can generate SR directly from H&E, yet systematic benchmarks of unsupervised models are scarce and their predictive uncertainty has not been quantified, even though visually plausible outputs may not faithfully reproduce the underlying tissue structure. We therefore benchmark six unsupervised image-to-image architectures (GAN-based and diffusion-based) across 54 scaling configurations on a newly released paired H&E to SR mouse liver dataset, the first open resource for this translation task. Each configuration is evaluated jointly on perceptual, distributional, and task-specific axes plus a blinded expert reader study; the best per family is then retrained as a deep ensemble, the first systematic comparison of epistemic uncertainty across unsupervised stain-to-stain architectures. Across families, perceptual quality, task-specific error, and ensemble agreement measure largely independent axes of model fitness: GAN-based methods cluster tightly on perceptual metrics yet differ substantially on task error and ensemble agreement, while the diffusion-based method (CycleDiffusion) is qualitatively different on all three. No single metric captures these differences, so reliable virtual staining requires reporting and selecting on all three jointly. The dataset, tiling pipeline, models, and evaluation code are released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。