arXiv:2605.25944cs.CVcs.AI2026-05中稿 · MICCAI 2026

无需训练,单点点击即可精准分割超声视频,抗噪声且不漂移。

EchoPilot: Training-Free Ultrasound Video Segmentation via Scale-Space Semantic Prompting and Reliability-Gated Memory

论文配图:EchoPilot: Training-Free Ultrasound Video Segmentation via Scale-Space Semantic Prompting and Reliability-Gated Memory
图 1 · 摘自论文原文
  • 用多尺度语义提示解决初始点信息不足问题
  • 动态门控记忆机制显著降低分割误差累积
  • 适合临床医生快速标注胎儿胎盘超声视频

超声视频分割在临床中极具价值,但受斑点噪声、边界模糊和快速解剖形变影响,难度极大。现有可提示基础模型虽支持点引导分割,但在超声场景下仍不可靠:单点提供的空间上下文不足,难以解决尺度歧义;贪婪式记忆更新会将早期错误放大为严重的时间漂移。本文提出EchoPilot,一种无需训练的超声视频分割框架,仅需第一帧单点点击和解剖类别名称。该框架协同冻结的医学视觉语言模型(VLM)进行语义定位、视觉基础模型(VFM)提取密集几何特征,并通过可提示视频分割器生成掩码并传播。为解决初始化歧义,提出尺度空间语义提示,先通过无参的S.E.E.D.(语义能量-熵密度)准则选择最优上下文视图,再从密集基础特征合成几何精确的辅助点提示,无需额外交互。为减少传播漂移,引入可靠性门控记忆更新机制,在预测不确定时选择性冻结分割器的记忆库,防止误差积累。我们还构建了首个动态胎儿胎盘超声视频数据集,含671帧标注。在三个超声视频数据集上,EchoPilot在稀疏交互设置下达到当前最佳性能,持续优于无训练基线与微调专用模型。

原文摘要 · Abstract (English)

Ultrasound video segmentation is clinically valuable yet difficult due to speckle noise, weak boundaries, and rapid anatomical deformation. Recent promptable foundation models enable point-guided segmentation, but their direct deployment in ultrasound remains unreliable: a single point provides insufficient spatial context to resolve scale ambiguity, and greedy memory updates amplify early errors into severe temporal drift. We present EchoPilot, a training-free framework for ultrasound video segmentation under sparse first-frame interaction, requiring only a single point click and an anatomical category name. EchoPilot orchestrates a frozen medical vision-language model (VLM) for semantic localization, a vision foundation model (VFM) for dense geometric feature extraction, and a promptable video segmentor for mask prediction and propagation. To resolve initialization ambiguity, we propose Scale-Space Semantic Prompting, which first selects an optimal contextual view via a parameter-free S.E.E.D. (Semantic Energy-Entropy Density) criterion, and then synthesizes geometrically precise auxiliary point prompts from dense foundation features without additional user interaction. To reduce propagation drift, a Reliability-Gated Memory update is further introduced to selectively freeze the segmentor's memory bank under uncertain predictions, preventing error accumulation. We also contribute the first dynamic fetal placenta ultrasound video segmentation dataset with 671 annotated frames. Across three ultrasound video datasets, EchoPilot achieves state-of-the-art performance under the sparse-interactive setting, consistently outperforming training-free baselines and finetuned specialists.

超声分割零样本视频分割医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。