arXiv:2503.20258cs.CVcs.AI2025-03被引 1

用3D结构建模提升超声视频分析精度与数据效率

Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos

  • 基于视觉马尔可夫网络构建3D视频模型,保留时空关联性
  • 在4个数据集上达到当前最佳性能,小样本下仍表现优异
  • 适合医疗影像分析、数据稀缺场景下的模型开发

超声视频是重要的临床影像数据,深度学习可提升诊断准确性和效率。但标注数据稀缺及视频分析固有挑战制约了方法发展。本文提出E-ViM³,一种保持视频3D结构的数据高效视觉马尔可夫网络,增强长程依赖与归纳偏置,更好建模时空相关性。通过引入包围式全局令牌(EGT),模型更有效捕捉和聚合全局特征。为提升数据效率,采用自监督预训练,设计适配多种视频场景的时空链式(STC)掩码策略。实验表明,E-ViM³在四个不同规模数据集(EchoNet-Dynamic、CAMUS、MICCAI-BUV、WHBUS)的两个高层语义分析任务中均达到当前最优表现。且在标签有限情况下仍具竞争力,展现出在真实临床应用中的潜力。

原文摘要 · Abstract (English)

Ultrasound videos are an important form of clinical imaging data, and deep learning-based automated analysis can improve diagnostic accuracy and clinical efficiency. However, the scarcity of labeled data and the inherent challenges of video analysis have impeded the advancement of related methods. In this work, we introduce E-ViM$^3$, a data-efficient Vision Mamba network that preserves the 3D structure of video data, enhancing long-range dependencies and inductive biases to better model space-time correlations. With our design of Enclosure Global Tokens (EGT), the model captures and aggregates global features more effectively than competing methods. To further improve data efficiency, we employ masked video modeling for self-supervised pre-training, with the proposed Spatial-Temporal Chained (STC) masking strategy designed to adapt to various video scenarios. Experiments demonstrate that E-ViM$^3$ performs as the state-of-the-art in two high-level semantic analysis tasks across four datasets of varying sizes: EchoNet-Dynamic, CAMUS, MICCAI-BUV, and WHBUS. Furthermore, our model achieves competitive performance with limited labels, highlighting its potential impact on real-world clinical applications.

医学影像视频分析自监督学习数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。