arXiv:2508.13077eess.IVcs.AI2025-08被引 2

用少量数据将胸超声模型转为食管超声生成,提升心脏结构分割效果。

From Transthoracic to Transesophageal: Cross-Modality Generation using LoRA Diffusion

  • 用低秩适配与轻量映射层,仅需少量新数据即可迁移模型。
  • 合成图像使多类分割的Dice分数提升,尤其改善右心结构识别。
  • 适合医疗影像数据少、需跨模态生成的研究者使用。

深度扩散模型在真实图像生成方面表现优异,但需要大量训练数据,在如经食管超声心动图(TEE)这类数据稀缺的领域面临挑战。尽管合成数据已提升经胸超声(TTE)性能,但TEE仍严重缺乏代表性,制约了深度学习在此高影响力模态中的应用。我们通过仅使用少量新病例和参数量小至 $10^5$ 的适配器,将预训练的掩码条件扩散模型从TTE迁移到TEE。该流程结合低秩适配(LoRA)与MaskR$^2$——一种轻量级重映射层,可将新掩码格式对齐至预训练模型的条件通道。此设计使用户能将模型适配到原始模型未覆盖的解剖结构集。通过定向适配策略,仅微调MLP层即足以生成高质量TEE图像。将少于200帧真实TEE图像与合成样本混合后,多类分割任务的Dice分数显著提升,尤其改善了被低估的右心结构识别。结果表明:(1) 可以低开销生成语义可控的TEE图像;(2) MaskR$^2$有效转换未知掩码格式且不损害下游任务性能;(3) 生成图像能有效提升多类分割任务表现。

原文摘要 · Abstract (English)

Deep diffusion models excel at realistic image synthesis but demand large training sets-an obstacle in data-scarce domains like transesophageal echocardiography (TEE). While synthetic augmentation has boosted performance in transthoracic echo (TTE), TEE remains critically underrepresented, limiting the reach of deep learning in this high-impact modality. We address this gap by adapting a TTE-trained, mask-conditioned diffusion backbone to TEE with only a limited number of new cases and adapters as small as $10^5$ parameters. Our pipeline combines Low-Rank Adaptation with MaskR$^2$, a lightweight remapping layer that aligns novel mask formats with the pretrained model's conditioning channels. This design lets users adapt models to new datasets with a different set of anatomical structures to the base model's original set. Through a targeted adaptation strategy, we find that adapting only MLP layers suffices for high-fidelity TEE synthesis. Finally, mixing less than 200 real TEE frames with our synthetic echoes improves the dice score on a multiclass segmentation task, particularly boosting performance on underrepresented right-heart structures. Our results demonstrate that (1) semantically controlled TEE images can be generated with low overhead, (2) MaskR$^2$ effectively transforms unseen mask formats into compatible formats without damaging downstream task performance, and (3) our method generates images that are effective for improving performance on a downstream task of multiclass segmentation.

医学图像生成跨模态低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。