arXiv:2606.31007cs.CV2026-06

用视觉模型先验实现手术视频中关键操作点的精准定位,无需人工标注。

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

  • 以SAM3为结构先验,通过点提示生成密集器械上下文。
  • 在无像素级标注下,尖端与锚点定位F1分别达72.4%和58.0%。
  • 粗到精的渐进式优化提升性能,适合手术动作分析场景。

视觉基础模型如SAM3可在多样化的手术视频条件下提供可迁移的物体级结构,但其分割结果未显式编码与操作相关的语义,定义功能性手术标志点。器械范围与几何估计不同于夹持、抓取或解剖相关的尖端或锚点定位。本文研究基于视觉基础模型的稀疏动作感知标志点定位,采用零样本、点提示的结构掩码,在无需人工像素级标注的前提下,提供密集器械级上下文。提出轻量级精炼框架,以SAM3作为结构先验:粗粒度多帧网络预测尖端与锚点提示,生成非真值掩码,融合视觉与热力图特征以精炼功能标志点预测。对比直接掩码增强监督、预测衍生掩码先验精炼及辅助掩码监督,探究视觉基础模型提供的结构如何融入高精度定位系统。在来自YouTube、Cholec80、HeiChole、SurgVU和CRCD的60段手术视频共7,867个片段上评估,该方法在异构条件下表现良好。无需手动像素级标注训练,所提模型在尖端定位上获得72.4%的整体F1分数,在锚点定位上达58.0%。消融实验表明,粗到精的精炼带来显著性能提升,且将预测衍生结构先验作为中间引导比作为直接定位目标更具优势。

原文摘要 · Abstract (English)

Vision foundation models such as SAM 3 can provide transferable object-level structure across diverse surgical video conditions, but segmentation outputs do not explicitly encode the action-conditioned semantics that define functional surgical landmarks. Estimating instrument extent and geometry differs from localizing the tip or anchor relevant to clipping, grasping, or dissecting. We investigate vision foundation model-enabled sparse action-aware landmark localization, using zero-shot, point-prompted structural masks to provide dense instrument-level context without manual pixel-level mask annotations. We propose a lightweight refinement framework that uses SAM 3 as a structural prior. A coarse multi-frame network predicts tip and anchor prompts, generating non-oracle masks that are fused with visual and heatmap features to refine functional landmark predictions. We compare direct mask-augmented supervision, prediction-derived mask-prior refinement, and auxiliary mask supervision to examine how vision foundation model-derived structure should enter a precision-oriented localization system. Experiments on 7,867 clips from 60 surgical videos spanning YouTube, Cholec80, HeiChole, SurgVU, and CRCD evaluate the approach under heterogeneous conditions. Without manual pixel-level mask annotations for training, the proposed model achieves overall F1 scores of 72.4% for tip and 58.0% for anchor localization. Ablations show that coarse-to-fine refinement provides a substantial performance gain, while prediction-derived structural priors provide additional improvement when incorporated as intermediate guidance rather than direct localization targets.

手术视频地标定位视觉模型无标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。