通过数据实验发现,超声视频分割中数据量与时间上下文比模型架构更重要。
Towards Better Ultrasound Video Segmentation Foundation Model: An Empirical study on SAM2 Finetuning from Data Perspective
- 从数据角度系统评估SAM2在超声视频上的微调效果,考察数据规模、时长和增强策略影响。
- 在三个数据集上验证:数据量大+时序信息充分时,性能显著提升,超越模型结构改进。
- 提出六种超声专用增强方法,联合训练可平衡模态对齐与任务特异性,适合医疗场景应用。
超声视频分割因数据集间与集内差异大、运动伪影多及标注数据有限而极具挑战。尽管基础模型如分割一切模型2(SAM2)具备强大的零样本与提示引导分割能力,但其迁移到医学影像领域时性能明显下降。现有适应研究多聚焦于结构改进,而对数据特征与训练范式的影响尚未系统探讨。本研究开展以数据为中心的SAM2适应性全面分析,考察训练集规模、视频时长与增强策略在三种范式(任务特定微调、中间适应、多任务联合训练)下的影响,涵盖五种SAM2变体与多种提示方式。进一步设计六种超声专用增强方法,并与通用策略对比。在三个代表性超声数据集上的实验表明,数据规模与时间上下文对适应性能起决定性作用,优于模型架构或初始化。此外,联合训练在模态对齐与任务专精之间提供了高效折中方案。本工作旨在为构建高效、数据敏感的SAM2超声视频分析适配流程提供实证依据。
原文摘要 · Abstract (English)
Ultrasound (US) video segmentation remains a challenging problem due to strong inter- and intra-dataset variability, motion artifacts, and limited annotated data. Although foundation models such as Segment Anything Model 2 (SAM2) demonstrate strong zero-shot and prompt-guided segmentation capabilities, their performance deteriorates substantially when transferred to medical imaging domains. Current adaptation studies mainly emphasize architectural modifications, while the influence of data characteristics and training regimes has not been systematically examined. In this study, we present a comprehensive, data-centric investigation of SAM2 adaptation for ultrasound video segmentation. We analyze how training-set size, video duration, and augmentation schemes affect adaptation performance under three paradigms: task-specific fine-tuning, intermediate adaptation, and multi-task joint training, across five SAM2 variants and multiple prompting modes. We further design six ultrasound-specific augmentations, assessing their effect relative to generic strategies. Experiments on three representative ultrasound datasets reveal that data scale and temporal context play a more decisive role than model architecture or initialization. Moreover, joint training offers an efficient compromise between modality alignment and task specialization. This work aims to provide empirical insights for developing efficient, data-aware adaptation pipelines for SAM2 in ultrasound video analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。