用语义与真实组织切片联合指导,生成高保真病理图像。
Semantic and Visual Crop-Guided Diffusion Models for Heterogeneous Tissue Synthesis in Histopathology
- 结合语义图与真实组织块,直接注入关键形态细节。
- 在Camelyon16上生成图像的弗雷谢距离降低6倍,达72.0。
- 无需人工标注即可处理11,765张全切片图像,适合数据稀缺场景。
病理图像合成面临独特挑战:保持组织异质性、捕捉细微形态特征,并扩展至未标注数据集。我们提出一种潜在扩散模型,通过结合语义分割图与组织特异性视觉切片的双重条件机制,生成逼真的异质性病理图像。对于标注数据集(如Camelyon16、Panda),提取的切片确保组织异质性在20%-80%之间;对于未标注数据(如TCGA),引入自监督扩展,利用基础模型嵌入将全切片图像聚类为100种组织类型,自动生成伪语义图用于训练。该方法生成高保真图像并附带精确区域标注,在下游分割任务中表现优异。在标注数据集上,仅用合成数据训练的模型性能接近真实数据基线,证明可控异质组织生成的有效性。定量评估显示,提示引导合成使Camelyon16的弗雷谢距离降低6倍(从430.1降至72.0),在Panda和TCGA上降低2-3倍。基于合成数据训练的DeepLabv3+模型在Camelyon16和Panda上的测试交并比分别为0.71和0.95,仅比真实数据基线低1-2%。该框架可扩展至11,765张TCGA全切片图像,无需人工标注,为计算病理学中多样且带注释的数据生成提供实用解决方案,缓解关键瓶颈。
原文摘要 · Abstract (English)
Synthetic data generation in histopathology faces unique challenges: preserving tissue heterogeneity, capturing subtle morphological features, and scaling to unannotated datasets. We present a latent diffusion model that generates realistic heterogeneous histopathology images through a novel dual-conditioning approach combining semantic segmentation maps with tissue-specific visual crops. Unlike existing methods that rely on text prompts or abstract visual embeddings, our approach preserves critical morphological details by directly incorporating raw tissue crops from corresponding semantic regions. For annotated datasets (i.e., Camelyon16, Panda), we extract patches ensuring 20-80% tissue heterogeneity. For unannotated data (i.e., TCGA), we introduce a self-supervised extension that clusters whole-slide images into 100 tissue types using foundation model embeddings, automatically generating pseudo-semantic maps for training. Our method synthesizes high-fidelity images with precise region-wise annotations, achieving superior performance on downstream segmentation tasks. When evaluated on annotated datasets, models trained on our synthetic data show competitive performance to those trained on real data, demonstrating the utility of controlled heterogeneous tissue generation. In quantitative evaluation, prompt-guided synthesis reduces Frechet Distance by up to 6X on Camelyon16 (from 430.1 to 72.0) and yields 2-3x lower FD across Panda and TCGA. Downstream DeepLabv3+ models trained solely on synthetic data attain test IoU of 0.71 and 0.95 on Camelyon16 and Panda, within 1-2% of real-data baselines (0.72 and 0.96). By scaling to 11,765 TCGA whole-slide images without manual annotations, our framework offers a practical solution for an urgent need for generating diverse, annotated histopathology data, addressing a critical bottleneck in computational pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。