用可控生成技术合成心脏瘢痕图像,提升医学影像分割精度
LGESynthNet: Controlled Scar Synthesis for Improved Scar Segmentation in Cardiac LGE-MRI Imaging
- 基于潜空间扩散模型,通过控制网实现对瘢痕位置、大小和深度的精准控制
- 仅用429张图像训练,合成样本使分割准确率最高提升20个百分点
- 适合需要小样本医学数据增强的研究者和医疗AI开发者
LGE心脏MRI中增强区域的分割对诊断缺血性和非缺血性心肌病至关重要。然而,像素级标注困难且耗时,导致标注数据稀缺。生成模型尤其是扩散模型在合成数据方面具有潜力,但多数依赖大规模训练数据,且难以精细控制小而局部的特征。我们提出LGESynthNet,一种基于潜在扩散的可控增强合成框架,可明确控制瘢痕的大小、位置和透壁程度。该方法采用基于ControlNet的修复架构,集成:(a) 针对条件监督的奖励模型,(b) 提供解剖描述性文本提示的标注模块,(c) 生物医学文本编码器。仅在429张图像(79名患者)上训练,即可生成解剖一致的逼真样本。通过质量控制过滤器筛选出高条件保真度输出,用于训练增强后,下游分割与检测性能分别提升最多6和20个百分点。
原文摘要 · Abstract (English)
Segmentation of enhancement in LGE cardiac MRI is critical for diagnosing various ischemic and non-ischemic cardiomyopathies. However, creating pixel-level annotations for these images is challenging and labor-intensive, leading to limited availability of annotated data. Generative models, particularly diffusion models, offer promise for synthetic data generation, yet many rely on large training datasets and often struggle with fine-grained conditioning control, especially for small or localized features. We introduce LGESynthNet, a latent diffusion-based framework for controllable enhancement synthesis, enabling explicit control over size, location, and transmural extent. Formulated as inpainting using a ControlNet-based architecture, the model integrates: (a) a reward model for conditioning-specific supervision, (b) a captioning module for anatomically descriptive text prompts, and (c) a biomedical text encoder. Trained on just 429 images (79 patients), it produces realistic, anatomically coherent samples. A quality control filter selects outputs with high conditioning-fidelity, which when used for training augmentation, improve downstream segmentation and detection performance, by up-to 6 and 20 points respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。