arXiv:2506.23440cs.CV2025-06ICCV被引 8

用无配对文本和掩码生成病理图像,提升细节与语义控制。

PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions

  • 将文本与掩码映射到统一条件空间,实现无配对数据联合建模。
  • 生成图像在结构精度与文本一致性上优于现有方法。
  • 适合需要高保真病理图像增强的医学影像研究者。

基于扩散的生成模型在缓解因隐私限制导致的数据稀缺问题方面展现出潜力。诊断文本报告提供高层次语义描述,掩码则包含区分形态区域的精细空间结构。然而,公开数据集缺乏同一病理图像的配对文本与掩码,限制了二者在图像生成中的联合使用。为此,我们提出PathDiff,一种通过将文本与掩码融合至统一条件空间,从无配对数据中有效学习的扩散框架。该方法可精准控制结构与上下文特征,生成高质量、语义准确的病理图像,并在图像保真度、文本-图像对齐及忠实度方面均有所提升,显著增强下游任务(如核分割与分类)的数据增广效果。大量实验验证其优于现有方法。

原文摘要 · Abstract (English)

Diffusion-based generative models have shown promise in synthesizing histopathology images to address data scarcity caused by privacy constraints. Diagnostic text reports provide high-level semantic descriptions, and masks offer fine-grained spatial structures essential for representing distinct morphological regions. However, public datasets lack paired text and mask data for the same histopathological images, limiting their joint use in image generation. This constraint restricts the ability to fully exploit the benefits of combining both modalities for enhanced control over semantics and spatial details. To overcome this, we propose PathDiff, a diffusion framework that effectively learns from unpaired mask-text data by integrating both modalities into a unified conditioning space. PathDiff allows precise control over structural and contextual features, generating high-quality, semantically accurate images. PathDiff also improves image fidelity, text-image alignment, and faithfulness, enhancing data augmentation for downstream tasks like nuclei segmentation and classification. Extensive experiments demonstrate its superiority over existing methods.

病理图像生成扩散模型多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。