PixCell是首个病理图像生成基础模型,可合成高保真病理图用于数据增强与隐私保护。
PixCell: A generative foundation model for digital histopathology images
- 基于扩散模型,无标注数据训练,利用自监督条件生成
- 在小样本数据上提升分类性能,生成图像与真实数据特征一致
- 支持隐私保护数据共享与虚拟免疫组化染色,适合医疗研究者使用
组织学切片的数字化推动了癌症诊断与研究的发展,形成了大规模数据集。自监督与视觉-语言模型已被证明能有效挖掘病理数据以学习判别性表征。然而,病理领域存在标注数据稀缺、数据共享受隐私法规限制,以及虚拟染色等固有的生成任务挑战。生成模型通过图像合成提供了解决方案。我们提出PixCell,首个面向病理图像的生成基础模型。PixCell是一种扩散模型,基于包含69,184张不同癌种H&E染色全切片图像的PanCan-30M大规模多样化数据集进行训练。采用渐进式训练策略与基于自监督的条件机制,实现无需人工标注即可扩展训练。通过真实切片作为条件,合成图像保留真实数据特性,可用于小规模数据集的数据增强以提升分类性能。我们通过两个生成下游任务验证PixCell的基础通用性:隐私保护的合成数据生成与虚拟免疫组化(IHC)染色。其高保真条件生成能力使机构能用私有数据生成高度逼真的、具有站点特异性的替代图像,替代原始患者数据共享。此外,利用约成对的H&E-IHC图像块,我们学习将条件从H&E迁移至多种IHC染色,实现从H&E输入生成IHC图像。训练模型已公开发布,以加速计算病理学研究。
原文摘要 · Abstract (English)
The digitization of histology slides has revolutionized pathology, providing massive datasets for cancer diagnosis and research. Self-supervised and vision-language models have been shown to effectively mine large pathology datasets to learn discriminative representations. On the other hand, there are unique problems in pathology, such as annotated data scarcity, privacy regulations in data sharing, and inherently generative tasks like virtual staining. Generative models, capable of synthesizing realistic and diverse images, present a compelling solution to address these problems through image synthesis. We introduce PixCell, the first generative foundation model for histopathology images. PixCell is a diffusion model trained on PanCan-30M, a large, diverse dataset derived from 69,184 H&E-stained whole slide images of various cancer types. We employ a progressive training strategy and a self-supervision-based conditioning that allows us to scale up training without any human-annotated data. By conditioning on real slides, the synthetic images capture the properties of the real data and can be used as data augmentation for small-scale datasets to boost classification performance. We prove the foundational versatility of PixCell by applying it to two generative downstream tasks: privacy-preserving synthetic data generation and virtual IHC staining. PixCell's high-fidelity conditional generation enables institutions to use their private data to synthesize highly realistic, site-specific surrogate images that can be shared in place of raw patient data. Furthermore, using datasets of roughly paired H&E-IHC tiles, we learn to translate PixCell's conditioning from H&E to multiple IHC stains, allowing the generation of IHC images from H&E inputs. Our trained models are publicly released to accelerate research in computational pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。