用零样本学习和双目视觉,实现无需训练的手术器械姿态精准估计。
SurgPose: Generalisable Surgical Instrument Pose Estimation using Zero-Shot Learning and Stereo Vision
- 结合零样本模型与双目深度估计,提升复杂环境下的姿态识别能力。
- 新方法在未见器械上表现优于现有模型,误差降低17.3%。
- 适合需要快速适配新器械的机器人手术系统开发者使用。
在机器人辅助微创手术(RMIS)中,手术器械的精确姿态估计对导航与机器人控制至关重要。传统标记法虽精度高,但易受遮挡、反光和器械设计限制;监督学习方法需大量标注数据,难以泛化至新器械。尽管零样本姿态估计在其他领域取得进展,但在RMIS中仍鲜有应用,导致对未见器械的泛化能力不足。本文提出一种6自由度(DoF)姿态估计新框架,融合前沿零样本RGB-D模型FoundationPose与SAM-6D。通过引入基于RAFT-Stereo的视觉深度估计,增强在反光和无纹理环境中的深度感知能力;同时,将SAM-6D中的实例分割模块替换为微调后的Mask R-CNN,显著提升遮挡及复杂场景下的分割精度。大量验证表明,改进后的SAM-6D在未见手术器械的零样本姿态估计中超越FoundationPose,创下新基准。该工作提升了对未知物体的姿态估计泛化能力,并首次将RGB-D零样本方法成功应用于RMIS。
原文摘要 · Abstract (English)
Accurate pose estimation of surgical tools in Robot-assisted Minimally Invasive Surgery (RMIS) is essential for surgical navigation and robot control. While traditional marker-based methods offer accuracy, they face challenges with occlusions, reflections, and tool-specific designs. Similarly, supervised learning methods require extensive training on annotated datasets, limiting their adaptability to new tools. Despite their success in other domains, zero-shot pose estimation models remain unexplored in RMIS for pose estimation of surgical instruments, creating a gap in generalising to unseen surgical tools. This paper presents a novel 6 Degrees of Freedom (DoF) pose estimation pipeline for surgical instruments, leveraging state-of-the-art zero-shot RGB-D models like the FoundationPose and SAM-6D. We advanced these models by incorporating vision-based depth estimation using the RAFT-Stereo method, for robust depth estimation in reflective and textureless environments. Additionally, we enhanced SAM-6D by replacing its instance segmentation module, Segment Anything Model (SAM), with a fine-tuned Mask R-CNN, significantly boosting segmentation accuracy in occluded and complex conditions. Extensive validation reveals that our enhanced SAM-6D surpasses FoundationPose in zero-shot pose estimation of unseen surgical instruments, setting a new benchmark for zero-shot RGB-D pose estimation in RMIS. This work enhances the generalisability of pose estimation for unseen objects and pioneers the application of RGB-D zero-shot methods in RMIS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。