StructDiff通过结构引导生成3D医学图像,精准还原复杂解剖细节。
StructDiff: Structure-aware Diffusion Model for 3D Fine-grained Medical Image Synthesis
- 引入图像-掩码配对模板,强化生成过程中的结构约束。
- 在公开数据集上,生成图像拓扑一致性与视觉真实感均达领先水平。
- 适合需要高精度解剖细节的医学图像生成与下游分割任务。
通过语义图像生成解决医学影像数据稀缺问题近年来备受关注。然而,现有生成模型主要聚焦于整体器官或大尺度组织结构的合成,难以复现细粒度解剖细节。由于对拓扑一致性要求严格且医学数据具有复杂的三维形态异质性,准确重建细粒度解剖细节仍是重大挑战。为此,我们提出 StructDiff——一种面向细粒度3D医学图像合成的结构感知扩散模型,可实现复杂解剖结构的精确生成。除传统的基于掩码的引导外,StructDiff进一步引入图像-掩码配对模板,提供结构约束并明确掩码与图像间的对应关系。同时,设计掩码生成模块(MGM)以丰富掩码多样性,缓解高质量参考掩码稀缺问题。此外,提出基于跳过采样方差(SSV)的置信度自适应学习(CAL)策略,降低不完美合成数据在下游任务迁移时带来的不确定性。大量实验表明,StructDiff在拓扑一致性和视觉真实性方面均达到当前最优性能,并显著提升下游分割表现。代码将在接受后发布。
原文摘要 · Abstract (English)
Solving medical imaging data scarcity through semantic image generation has attracted growing attention in recent years. However, existing generative models mainly focus on synthesizing whole-organ or large-tissue structures, showing limited capability in reproducing fine-grained anatomical details. Due to the stringent requirement of topological consistency and the complex 3D morphological heterogeneity of medical data, accurately reconstructing fine-grained anatomical details remains a significant challenge. To address these limitations, we propose StructDiff, a Structure-aware Diffusion Model for fine-grained 3D medical image synthesis, which enables precise generation of topologically complex anatomies. In addition to the conventional mask-based guidance, StructDiff further introduces a paired image-mask template to guide the generation process, providing structural constrains and offering explicit knowledge of mask-to-image correspondence. Moreover, a Mask Generation Module (MGM) is designed to enrich mask diversity and alleviate the scarcity of high-quality reference masks. Furthermore, we propose a Confidence-aware Adaptive Learning (CAL) strategy based on Skip-Sampling Variance (SSV), which mitigates uncertainty introduced by imperfect synthetic data when transferring to downstream tasks. Extensive experiments demonstrate that StructDiff achieves state-of-the-art performance in terms of topological consistency and visual realism, and significantly boosts downstream segmentation performance. Code will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。