用草图+文字生成结构准确的3D医学影像,解决数据少难题
Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generation
- 结合2D草图与文本描述,用扩散模型生成3D器官分割图
- 在公开数据集上生成体积与真实CT高度一致,结构准确率显著提升
- 适合医学图像生成、数据增强,尤其擅长可控低资源场景
扩散概率模型在生成高质量、逼真医学图像方面展现出巨大潜力,为医学领域长期存在的数据稀缺问题提供了有前景的解决方案。然而,在多模态条件下生成具有解剖学一致性结构的3D医学体数据仍是复杂且未解决的问题。我们提出Sketch2CT,一种基于多模态扩散的结构感知3D医学体生成框架,通过用户提供的2D草图和描述3D几何语义的文本联合引导。该框架首先从随机噪声生成目标器官的3D分割掩码,条件依赖于两种模态。为有效对齐和融合输入,我们设计两个关键模块:利用局部文本线索优化草图特征,并整合全局草图-文本表示。基于胶囊注意力主干网络,这些模块发挥草图与文本的互补优势,生成解剖学准确的器官形状。合成的分割掩码随后指导潜在扩散模型生成3D CT体数据,实现与用户定义草图和描述一致的器官外观真实重建。在公开CT数据集上的大量实验表明,Sketch2CT在生成多模态医学体数据方面表现卓越。其可控、低成本的生成流程支持医学数据集的规范高效扩充。代码已开源:https://github.com/adlsn/Sketch2CT。
原文摘要 · Abstract (English)
Diffusion probabilistic models have demonstrated significant potential in generating high-quality, realistic medical images, providing a promising solution to the persistent challenge of data scarcity in the medical field. Nevertheless, producing 3D medical volumes with anatomically consistent structures under multimodal conditions remains a complex and unresolved problem. We introduce Sketch2CT, a multimodal diffusion framework for structure-aware 3D medical volume generation, jointly guided by a user-provided 2D sketch and a textual description that captures 3D geometric semantics. The framework initially generates 3D segmentation masks of the target organ from random noise, conditioned on both modalities. To effectively align and fuse these inputs, we propose two key modules that refine sketch features with localized textual cues and integrate global sketch-text representations. Built upon a capsule-attention backbone, these modules leverage the complementary strengths of sketches and text to produce anatomically accurate organ shapes. The synthesized segmentation masks subsequently guide a latent diffusion model for 3D CT volume synthesis, enabling realistic reconstruction of organ appearances that are consistent with user-defined sketches and descriptions. Extensive experiments on public CT datasets demonstrate that Sketch2CT achieves superior performance in generating multimodal medical volumes. Its controllable, low-cost generation pipeline enables principled, efficient augmentation of medical datasets. Code is available at https://github.com/adlsn/Sketch2CT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。