arXiv:2503.22394cs.CVcs.AI2025-03

端镜下组织点追踪新框架,抗形变遮挡能力强。

Endo-TTAP: Robust Endoscopic Tissue Tracking via Multi-Facet Guided Attention and Hybrid Flow-point Supervision

  • 多模态注意力融合流场、语义特征与运动模式,预测带不确定性和遮挡感知
  • 两阶段课程学习,结合合成数据与伪标签提升长时追踪精度
  • 适用于复杂内镜场景,适合手术机器人导航研究者

内镜视频中精确的组织点追踪对机器人辅助手术导航和场景理解至关重要,但因复杂形变、器械遮挡及密集轨迹标注稀缺而极具挑战。现有方法在长期追踪中表现受限,主要源于特征利用不足与标注依赖。本文提出Endo-TTAP框架:(1) 多维度引导注意力(MFGA)模块,融合多尺度流场动态、DINOv2语义嵌入与显式运动模式,联合预测点位置并具备不确定性与遮挡感知能力;(2) 两阶段课程学习策略,采用辅助课程适配器(ACA)实现渐进初始化,并引入混合监督机制:第一阶段使用带光流真值的合成数据进行不确定性-遮挡正则化,第二阶段结合无监督流一致性与半监督学习,利用现成追踪器生成精炼伪标签。在两个MICCAI挑战数据集及自建数据集上的大量验证表明,该方法在复杂内镜条件下实现最优追踪性能。源代码与数据集将公开于https://anonymous.4open.science/r/Endo-TTAP-36E5。

原文摘要 · Abstract (English)

Accurate tissue point tracking in endoscopic videos is critical for robotic-assisted surgical navigation and scene understanding, but remains challenging due to complex deformations, instrument occlusion, and the scarcity of dense trajectory annotations. Existing methods struggle with long-term tracking under these conditions due to limited feature utilization and annotation dependence. We present Endo-TTAP, a novel framework addressing these challenges through: (1) A Multi-Facet Guided Attention (MFGA) module that synergizes multi-scale flow dynamics, DINOv2 semantic embeddings, and explicit motion patterns to jointly predict point positions with uncertainty and occlusion awareness; (2) A two-stage curriculum learning strategy employing an Auxiliary Curriculum Adapter (ACA) for progressive initialization and hybrid supervision. Stage I utilizes synthetic data with optical flow ground truth for uncertainty-occlusion regularization, while Stage II combines unsupervised flow consistency and semi-supervised learning with refined pseudo-labels from off-the-shelf trackers. Extensive validation on two MICCAI Challenge datasets and our collected dataset demonstrates that Endo-TTAP achieves state-of-the-art performance in tissue point tracking, particularly in scenarios characterized by complex endoscopic conditions. The source code and dataset will be available at https://anonymous.4open.science/r/Endo-TTAP-36E5.

医学图像目标追踪内镜手术注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。