用生成模型把超声彩超转成灰度图,让数据更均衡。
Generative deep learning for foundational video translation in ultrasound
- 设计双网络生成模型,融合像素、对抗和感知损失,还原真实超声图像。
- 合成视频与真实视频的结构相似度达0.91±0.04,临床专家难分辨。
- 模型跨领域适用,适用于心脏外多个临床场景,具基础性应用潜力。
深度学习有望革新医学影像采集与解读,但需关注数据不平衡与缺失问题。超声数据尤为复杂,除不同视角与结构外,还包含灰度和彩色多普勒(CFD)等多种子模态,临床研究中常出现数据不平衡。图像翻译可缓解此问题,但现有方法难以处理超声子模态。本文提出一种针对超声CFD-灰度视频转换的生成方法,基于54,975段视频训练,测试集为8,368段。该方法结合像素级、对抗性和感知损失,采用两个网络分别重建解剖结构与去噪,以实现逼真超声成像。合成视频与真实视频的平均成对SSIM为0.91±0.04。在深度学习分类与分割任务中,合成视频表现与真实视频无异;临床专家盲评显示:真实与合成视频的F1分数分别为0.9和0.89,分割Dice分数达0.97。总体判别准确率为54±6%(42–61%),表明合成视频高度逼真。尽管仅在心脏视频上训练,模型在多个临床领域仍表现良好(平均SSIM 0.91±0.05),展现基础能力。本研究拓展了回顾性影像数据的应用范围,并丰富了医学影像数据集设计工具箱。
原文摘要 · Abstract (English)
Deep learning (DL) has the potential to revolutionize image acquisition and interpretation across medicine, however, attention to data imbalance and missingness is required. Ultrasound data presents a particular challenge because in addition to different views and structures, it includes several sub-modalities-such as greyscale and color flow doppler (CFD)-that are often imbalanced in clinical studies. Image translation can help balance datasets but is challenging for ultrasound sub-modalities to date. Here, we present a generative method for ultrasound CFD-greyscale video translation, trained on 54,975 videos and tested on 8,368. The method developed leveraged pixel-wise, adversarial, and perceptual loses and utilized two networks: one for reconstructing anatomic structures and one for denoising to achieve realistic ultrasound imaging. Average pairwise SSIM between synthetic videos and ground truth was 0.91+/-0.04. Synthetic videos performed indistinguishably from real ones in DL classification and segmentation tasks and when evaluated by blinded clinical experts: F1 score was 0.9 for real and 0.89 for synthetic videos; Dice score between real and synthetic segmentation was 0.97. Overall clinician accuracy in distinguishing real vs synthetic videos was 54+/-6% (42-61%), indicating realistic synthetic videos. Although trained only on heart videos, the model worked well on ultrasound spanning several clinical domains (average SSIM 0.91+/-0.05), demonstrating foundational abilities. Together, these data expand the utility of retrospectively collected imaging and augment the dataset design toolbox for medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。