arXiv:2603.03125cs.CV2026-03中稿 · ICASSP 2026

用小波变换增强扩散模型,生成更逼真的肺部超声图像。

AWDiff: An a trous wavelet diffusion model for lung ultrasound image synthesis

  • 采用a trous小波变换保留细粒度结构,避免降采样损失
  • 在真实数据集上比现有方法更低失真、更高感知质量
  • 结合医学视觉语言模型,生成符合临床标签的图像

肺部超声(LUS)是一种安全且便携的影像技术,但数据稀缺限制了机器学习在图像解读与疾病监测中的应用。现有生成增强方法如生成对抗网络(GANs)和扩散模型常因分辨率降低而丢失细微诊断特征,尤其是B线和胸膜不规则。本文提出基于扩散模型的a trous小波扩散框架(AWDiff),通过a trous小波变换在生成过程中保持细尺度结构,避免破坏性下采样。同时,引入基于BioMedCLIP(一种在大规模生物医学语料上训练的视觉语言基础模型)的语义条件控制,确保生成图像与临床有意义标签对齐。在肺部超声数据集上,AWDiff相比现有方法展现出更低的失真和更高的感知质量,验证了其在结构保真度与临床多样性方面的优势。

原文摘要 · Abstract (English)

Lung ultrasound (LUS) is a safe and portable imaging modality, but the scarcity of data limits the development of machine learning methods for image interpretation and disease monitoring. Existing generative augmentation methods, such as Generative Adversarial Networks (GANs) and diffusion models, often lose subtle diagnostic cues due to resolution reduction, particularly B-lines and pleural irregularities. We propose A trous Wavelet Diffusion (AWDiff), a diffusion based augmentation framework that integrates the a trous wavelet transform to preserve fine-scale structures while avoiding destructive downsampling. In addition, semantic conditioning with BioMedCLIP, a vision language foundation model trained on large scale biomedical corpora, enforces alignment with clinically meaningful labels. On a LUS dataset, AWDiff achieved lower distortion and higher perceptual quality compared to existing methods, demonstrating both structural fidelity and clinical diversity.

图像生成扩散模型医学影像小波变换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。