arXiv:2503.17475cs.CVcs.AI2025-03

提升超声视频分析精度,兼顾局部细节与全局空间信息。

Spatiotemporal Learning with Context-aware Video Tubelets for Ultrasound Video Analysis

  • 用位置、尺寸和置信度增强管状体特征,保留全局空间上下文。
  • 引入预训练检测模型的区域对齐特征图,扩大感受野并降低计算量。
  • 仅0.4M参数,适合实时超声病理检测,适用于临床部署。

基于视频成像的病理自动检测算法需准确解析复杂的时空信息,整合多帧观测结果。现有先进方法通过分类视频子体积(管状体)实现,但常因仅关注检测感兴趣区内的局部区域而丢失全局空间上下文。本文提出一种轻量级管状体基础目标检测与视频分类框架,同时保持全局空间上下文与精细时空特征。为缓解上下文丢失问题,将管状体的位置、大小和置信度作为分类器输入;此外,采用预训练检测模型生成的区域对齐特征图,利用其学习到的特征表示扩大感受野并减少计算复杂度。该方法高效,时空管状体分类器仅含0.4M参数。我们在14,804个来自828名患者的超声视频上,对肺实变和胸腔积液的检测与分类任务进行五折交叉验证,结果表明本方法优于以往管状体基方法,且适合实时工作流程。

原文摘要 · Abstract (English)

Computer-aided pathology detection algorithms for video-based imaging modalities must accurately interpret complex spatiotemporal information by integrating findings across multiple frames. Current state-of-the-art methods operate by classifying on video sub-volumes (tubelets), but they often lose global spatial context by focusing only on local regions within detection ROIs. Here we propose a lightweight framework for tubelet-based object detection and video classification that preserves both global spatial context and fine spatiotemporal features. To address the loss of global context, we embed tubelet location, size, and confidence as inputs to the classifier. Additionally, we use ROI-aligned feature maps from a pre-trained detection model, leveraging learned feature representations to increase the receptive field and reduce computational complexity. Our method is efficient, with the spatiotemporal tubelet classifier comprising only 0.4M parameters. We apply our approach to detect and classify lung consolidation and pleural effusion in ultrasound videos. Five-fold cross-validation on 14,804 videos from 828 patients shows our method outperforms previous tubelet-based approaches and is suited for real-time workflows.

超声视频时空学习管状体轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。