arXiv:2410.23413cs.CV2024-10被引 25

EchoFM是首个用于超声心动图视频的通用基础模型,无需标签即可学习视频特征。

EchoFM: Foundation Model for Generalizable Echocardiogram Analysis

  • 采用时空一致性掩码与周期驱动对比学习,自监督捕捉视频动态
  • 在29万段超声心动图视频上预训练,覆盖26种扫描视角
  • 可适配多种下游任务,性能超越现有专用与通用模型

基础模型因其跨任务和数据分布的泛化能力而受到广泛关注。尽管医疗领域已出现基础模型,但针对心脏影像,尤其是超声心动图视频的解决方案仍属空白。本文提出EchoFM,一种专为超声心动图视频设计的基础模型。其采用自监督学习框架,通过时空一致性掩码策略与周期驱动对比学习,有效捕捉超声心动图的时空变化模式,并在无标签情况下学习代表性视频特征。模型在包含超过29万段超声心动图视频、覆盖26种扫描视角、总帧数达2000万的海量数据集上进行预训练。预训练后的EchoFM可轻松适配并微调至多种下游任务,作为强大骨干模型使用。系统评估涵盖四类检查后下游任务,实验结果表明,EchoFM在所有任务中均优于当前最先进方法,包括专用超声心动图模型、自监督预训练模型及通用预训练基础模型。

原文摘要 · Abstract (English)

Foundation models have recently gained significant attention because of their generalizability and adaptability across multiple tasks and data distributions. Although medical foundation models have emerged, solutions for cardiac imaging, especially echocardiography videos, are still unexplored. In this paper, we introduce EchoFM, a foundation model specifically designed to represent and analyze echocardiography videos. In EchoFM, we propose a self-supervised learning framework that captures both spatial and temporal variability patterns through a spatio-temporal consistent masking strategy and periodic-driven contrastive learning. This framework can effectively capture the spatio-temporal dynamics of echocardiography and learn the representative video features without any labels. We pre-train our model on an extensive dataset comprising over 290,000 echocardiography videos covering 26 scan views across different imaging modes, with up to 20 million frames of images. The pre-trained EchoFM can then be easily adapted and fine-tuned for a variety of downstream tasks, serving as a robust backbone model. Our evaluation was systemically designed for four downstream tasks after the echocardiography examination routine. Experiment results show that EchoFM surpasses state-of-the-art methods, including specialized echocardiography methods, self-supervised pre-training models, and general-purposed pre-trained foundation models, across all downstream tasks.

超声心动图基础模型自监督学习视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。