无需训练即可精确定位未见过的手术器械6D姿态
Training-free Detection and 6D Pose Estimation of Unseen Surgical Instruments
- 仅用纹理CAD模型,通过多视角匹配与几何一致性检测
- 毫米级精度,性能媲美有监督方法且可泛化到新器械
- 适合动态临床场景下的无标记器械追踪,无需标注数据
准确检测与6D姿态估计对计算机辅助手术至关重要。现有监督方法缺乏对新器械的灵活性且需大量标注数据。本文提出一种无需训练的多视角6D姿态估计流程,仅需带纹理的CAD模型作为先验知识。首先,通过预训练特征提取器在各视角生成对象掩码候选,并基于渲染模板相似性评分;跨视角匹配后三角化为3D实例候选,再通过多视角几何一致性过滤。其次,采用特征-度量评分与跨视角注意力迭代优化姿态假设,最终利用新型多视角、遮挡感知轮廓配准最小化未遮挡轮廓点的重投影误差。在真实手术数据集MVPSP上评估显示,该方法在控制条件下实现毫米级精度,与有监督方法相当,同时具备对未见器械的完全泛化能力。结果验证了无训练、无标记在手术场景中的可行性,并揭示了手术环境的独特挑战。本方法结合前沿基础模型、多视角几何与轮廓优化,实现无需任务特定训练的高精度器械姿态估计,适用于动态临床环境中的鲁棒器械追踪与场景理解。
原文摘要 · Abstract (English)
Purpose: Accurate detection and 6D pose estimation of surgical instruments are crucial for many computer-assisted interventions. However, supervised methods lack flexibility for new or unseen tools and require extensive annotated data. This work introduces a training-free pipeline for accurate multi-view 6D pose estimation of unseen surgical instruments, which only requires a textured CAD model as prior knowledge. Methods: Our pipeline consists of two main stages. First, for detection, we generate object mask proposals in each view and score their similarity to rendered templates using a pre-trained feature extractor. Detections are matched across views, triangulated into 3D instance candidates, and filtered using multi-view geometric consistency. Second, for pose estimation, a set of pose hypotheses is iteratively refined and scored using feature-metric scores with cross-view attention. The best hypothesis undergoes a final refinement using a novel multi-view, occlusion-aware contour registration, which minimizes reprojection errors of unoccluded contour points. Results: The proposed method was rigorously evaluated on real-world surgical data from the MVPSP dataset. The method achieves millimeter-accurate pose estimates that are on par with supervised methods under controlled conditions, while maintaining full generalization to unseen instruments. These results demonstrate the feasibility of training-free, marker-less detection and tracking in surgical scenes, and highlight the unique challenges in surgical environments. Conclusion: We present a novel and flexible pipeline that effectively combines state-of-the-art foundational models, multi-view geometry, and contour-based refinement for high-accuracy 6D pose estimation of surgical instruments without task-specific training. This approach enables robust instrument tracking and scene understanding in dynamic clinical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。