arXiv:2512.09418cs.CV2025-12中稿 · MICAD 2025

无需标注数据,用运动特征生成真实心动超声视频

Label-free Motion-Conditioned Diffusion Model for Cardiac Ultrasound Synthesis

  • 用自监督方式提取视频运动与外观特征,解耦表示
  • 在EchoNet-Dynamic上生成时序连贯、临床真实的超声视频
  • 适合医疗图像生成、数据稀缺场景的算法研究者

超声心动图对心脏功能的无创、实时评估至关重要,但受限于隐私保护和专家标注复杂性,标注数据稀少,严重制约深度学习方法发展。本文提出运动条件扩散模型(MCDM),一种无需标签的潜在扩散框架,可基于自监督运动特征生成逼真的心动超声视频。为提取特征,设计了运动与外观特征提取器(MAFE),从视频中解耦运动与外观表示。通过伪外观特征引导的重识别损失和伪光流场引导的光流损失,进一步增强特征学习。在EchoNet-Dynamic数据集上评估显示,MCDM实现了具有竞争力的视频生成性能,生成时序连贯且临床真实的序列,完全不依赖人工标注。结果表明,自监督条件化在可扩展心动超声合成中具有巨大潜力。代码已开源:https://github.com/ZheLi2020/LabelfreeMCDM。

原文摘要 · Abstract (English)

Ultrasound echocardiography is essential for the non-invasive, real-time assessment of cardiac function, but the scarcity of labelled data, driven by privacy restrictions and the complexity of expert annotation, remains a major obstacle for deep learning methods. We propose the Motion Conditioned Diffusion Model (MCDM), a label-free latent diffusion framework that synthesises realistic echocardiography videos conditioned on self-supervised motion features. To extract these features, we design the Motion and Appearance Feature Extractor (MAFE), which disentangles motion and appearance representations from videos. Feature learning is further enhanced by two auxiliary objectives: a re-identification loss guided by pseudo appearance features and an optical flow loss guided by pseudo flow fields. Evaluated on the EchoNet-Dynamic dataset, MCDM achieves competitive video generation performance, producing temporally coherent and clinically realistic sequences without reliance on manual labels. These results demonstrate the potential of self-supervised conditioning for scalable echocardiography synthesis. Our code is available at https://github.com/ZheLi2020/LabelfreeMCDM.

超声合成扩散模型自监督医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。