arXiv:2603.26834eess.IVcs.AI2026-03中稿 · IEEE International…

用混合扩散模型生成更逼真的乳腺超声图像,提升诊断建模质量。

Hybrid Diffusion Model for Breast Ultrasound Image Augmentation

  • 结合文生图与图生图细化,保留超声纹理特征。
  • FID从45.97降至33.29,图像保真度显著提升。
  • 适合医学图像增强与小样本医疗建模研究者使用。

我们提出一种基于混合扩散的数据增强框架,以应对乳腺超声(BUS)数据集中的关键挑战。与传统扩散增强不同,该方法通过结合文本到图像生成与图像到图像(img2img)精炼,并采用低秩适配(LoRA)和文本反转(TI)进行微调,提升了图像视觉保真度并保留了超声纹理特征。在开源的Kaggle乳腺超声图像数据集(BUSI)上,生成了真实且类别一致的图像。相较于Stable Diffusion v1.5基线,引入TI和img2img精炼后,弗雷切特起始距离(FID)从45.97降至33.29,表明保真度大幅提升,同时下游分类性能保持相当。整体而言,该框架有效缓解了合成超声图像保真度不足的问题,增强了增强效果,为鲁棒诊断建模提供了支持。

原文摘要 · Abstract (English)

We propose a hybrid diffusion-based augmentation framework to overcome the critical challenge of ultrasound data augmentation in breast ultrasound (BUS) datasets. Unlike conventional diffusion-based augmentations, our approach improves visual fidelity and preserves ultrasound texture by combining text-to-image generation with image-to-image (img2img) refinement, as well as fine-tuning with low-rank adaptation (LoRA) and textual inversion (TI). Our method generated realistic, class-consistent images on an open-source Kaggle breast ultrasound image dataset (BUSI). Compared to the Stable Diffusion v1.5 baseline, incorporating TI and img2img refinement reduced the Frechet Inception Distance (FID) from 45.97 to 33.29, demonstrating a substantial gain in fidelity while maintaining comparable downstream classification performance. Overall, the proposed framework effectively mitigates the low-fidelity limitations of synthetic ultrasound images and enhances the quality of augmentation for robust diagnostic modeling.

超声图像扩散模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。