用扩散模型生成保留组织结构的合成病理图像,提升跨机构诊断泛化能力。
Semi-Supervised Domain Adaptation with Latent Diffusion for Pathology Image Classification
- 用源域和目标域无标签数据训练潜空间扩散模型,生成保形且适配目标域的合成图像。
- 在肺癌预后预测任务中,目标域测试集加权F1从0.611升至0.706,宏F1从0.641升至0.716。
- 适合需要跨医院泛化的病理图像分类场景,避免传统图像翻译导致的结构失真。
计算病理学中的深度学习模型常因域偏移难以在不同队列和机构间泛化。现有方法或无法利用目标域未标记数据,或依赖图像到图像转换,可能扭曲组织结构并降低模型精度。本文提出一种半监督域适应(SSDA)框架,利用在源域与目标域无标签数据上训练的潜空间扩散模型,生成保持形态且适配目标域的合成图像。通过条件控制基础模型特征、队列身份和组织制备方法,既保留源域组织结构,又引入目标域外观特征。将这些目标域感知的合成图像与源队列的真实标注图像联合训练下游分类器,并在目标队列上测试。在肺腺癌预后预测任务中验证了该框架的有效性:在不损害源队列性能的前提下,显著提升了目标队列测试集表现,加权F1从0.611增至0.706,宏F1从0.641增至0.716。结果表明,基于目标感知扩散的合成数据增强是提升计算病理学域泛化能力的有力方法。
原文摘要 · Abstract (English)
Deep learning models in computational pathology often fail to generalize across cohorts and institutions due to domain shift. Existing approaches either fail to leverage unlabeled data from the target domain or rely on image-to-image translation, which can distort tissue structures and compromise model accuracy. In this work, we propose a semi-supervised domain adaptation (SSDA) framework that utilizes a latent diffusion model trained on unlabeled data from both the source and target domains to generate morphology-preserving and target-aware synthetic images. By conditioning the diffusion model on foundation model features, cohort identity, and tissue preparation method, we preserve tissue structure in the source domain while introducing target-domain appearance characteristics. The target-aware synthetic images, combined with real, labeled images from the source cohort, are subsequently used to train a downstream classifier, which is then tested on the target cohort. The effectiveness of the proposed SSDA framework is demonstrated on the task of lung adenocarcinoma prognostication. The proposed augmentation yielded substantially better performance on the held-out test set from the target cohort, without degrading source-cohort performance. The approach improved the weighted F1 score on the target-cohort held-out test set from 0.611 to 0.706 and the macro F1 score from 0.641 to 0.716. Our results demonstrate that target-aware diffusion-based synthetic data augmentation provides a promising and effective approach for improving domain generalization in computational pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。