arXiv:2603.25228cs.CV2026-03中稿 · IJCARS: IPCAI 2026

无需训练即可精确定位未见过的手术器械6D姿态

Training-free Detection and 6D Pose Estimation of Unseen Surgical Instruments

  • 仅用纹理CAD模型,通过多视角匹配与几何一致性检测
  • 毫米级精度,性能媲美有监督方法且可泛化到新器械
  • 适合动态临床场景下的无标记器械追踪,无需标注数据

准确检测与6D姿态估计对计算机辅助手术至关重要。现有监督方法缺乏对新器械的灵活性且需大量标注数据。本文提出一种无需训练的多视角6D姿态估计流程,仅需带纹理的CAD模型作为先验知识。首先,通过预训练特征提取器在各视角生成对象掩码候选,并基于渲染模板相似性评分;跨视角匹配后三角化为3D实例候选,再通过多视角几何一致性过滤。其次,采用特征-度量评分与跨视角注意力迭代优化姿态假设,最终利用新型多视角、遮挡感知轮廓配准最小化未遮挡轮廓点的重投影误差。在真实手术数据集MVPSP上评估显示,该方法在控制条件下实现毫米级精度,与有监督方法相当,同时具备对未见器械的完全泛化能力。结果验证了无训练、无标记在手术场景中的可行性,并揭示了手术环境的独特挑战。本方法结合前沿基础模型、多视角几何与轮廓优化,实现无需任务特定训练的高精度器械姿态估计,适用于动态临床环境中的鲁棒器械追踪与场景理解。

原文摘要 · Abstract (English)

Purpose: Accurate detection and 6D pose estimation of surgical instruments are crucial for many computer-assisted interventions. However, supervised methods lack flexibility for new or unseen tools and require extensive annotated data. This work introduces a training-free pipeline for accurate multi-view 6D pose estimation of unseen surgical instruments, which only requires a textured CAD model as prior knowledge. Methods: Our pipeline consists of two main stages. First, for detection, we generate object mask proposals in each view and score their similarity to rendered templates using a pre-trained feature extractor. Detections are matched across views, triangulated into 3D instance candidates, and filtered using multi-view geometric consistency. Second, for pose estimation, a set of pose hypotheses is iteratively refined and scored using feature-metric scores with cross-view attention. The best hypothesis undergoes a final refinement using a novel multi-view, occlusion-aware contour registration, which minimizes reprojection errors of unoccluded contour points. Results: The proposed method was rigorously evaluated on real-world surgical data from the MVPSP dataset. The method achieves millimeter-accurate pose estimates that are on par with supervised methods under controlled conditions, while maintaining full generalization to unseen instruments. These results demonstrate the feasibility of training-free, marker-less detection and tracking in surgical scenes, and highlight the unique challenges in surgical environments. Conclusion: We present a novel and flexible pipeline that effectively combines state-of-the-art foundational models, multi-view geometry, and contour-based refinement for high-accuracy 6D pose estimation of surgical instruments without task-specific training. This approach enables robust instrument tracking and scene understanding in dynamic clinical environments.

6D姿态估计手术导航无监督方法多视角几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。