用多模态扩散模型生成精准病理图像,解决罕见细胞标注数据少的问题。
MSDM: Generating Task-Specific Pathology Images with a Multimodal Conditioned Diffusion Model for Cell and Nuclei Segmentation
- 通过形态、颜色和文本信息联合控制生成图像与掩码
- 合成数据使柱状细胞分割准确率显著提升
- 适合需要增强数据多样性的病理图像分割研究者
计算病理学中细胞和细胞核分割面临标注数据稀缺问题,尤其在罕见或非典型形态情况下。人工标注耗时费力,而合成数据可提供低成本替代方案。本文提出多模态语义扩散模型(MSDM),用于生成像素级精确的图像-掩码对。通过将细胞/核形态(水平与垂直图)、RGB颜色特征及BERT编码的实验/表型元数据作为条件,实现对生成图像的细粒度控制。各模态通过多头交叉注意力融合,确保生成结果符合特定生物学条件。定量分析显示,合成图像与真实图像在嵌入空间中的Wasserstein距离低,表明分布接近。以柱状细胞为例,引入合成样本后,分割模型性能显著提升。该方法系统性补充数据集,直接应对模型短板。验证了多模态扩散生成在提升细胞核分割模型鲁棒性与泛化能力方面的有效性,为生成模型在计算病理学中的广泛应用铺平道路。
原文摘要 · Abstract (English)
Scarcity of annotated data, particularly for rare or atypical morphologies, present significant challenges for cell and nuclei segmentation in computational pathology. While manual annotation is labor-intensive and costly, synthetic data offers a cost-effective alternative. We introduce a Multimodal Semantic Diffusion Model (MSDM) for generating realistic pixel-precise image-mask pairs for cell and nuclei segmentation. By conditioning the generative process with cellular/nuclear morphologies (using horizontal and vertical maps), RGB color characteristics, and BERT-encoded assay/indication metadata, MSDM generates datasests with desired morphological properties. These heterogeneous modalities are integrated via multi-head cross-attention, enabling fine-grained control over the generated images. Quantitative analysis demonstrates that synthetic images closely match real data, with low Wasserstein distances between embeddings of generated and real images under matching biological conditions. The incorporation of these synthetic samples, exemplified by columnar cells, significantly improves segmentation model accuracy on columnar cells. This strategy systematically enriches data sets, directly targeting model deficiencies. We highlight the effectiveness of multimodal diffusion-based augmentation for advancing the robustness and generalizability of cell and nuclei segmentation models. Thereby, we pave the way for broader application of generative models in computational pathology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。