arXiv:2509.13541cs.ROcs.CV2025-09

将语义分割与实时建图结合,提升内窥镜手术可视化与定位精度

PERSEUS: Perception with Semantic Endoscopic Understanding and SLAM

  • 融合学习式分割、深度估计与单目建图,实现手术场景实时语义地图
  • 重建误差低于1毫米,位姿误差0.9毫米,尺度估计误差小于2%
  • 适合需要高精度导航的微创机器人手术系统,助力术中自动化

目的:自然腔道手术相比开放手术创伤更小、恢复更快,但因视觉与定位困难,对医生要求更高。本文提出一种面向此类手术的感知流水线,实现场景语义理解。方法:将基于学习的分割、深度估计与3D重建模块集成,生成手术场景的实时语义地图;利用机器人位姿进行配准,解决单目图像的尺度模糊问题,支持在机器人手术中使用语义引导的实时重建。结果:基于平均单侧Chamfer距离的重建精度达到亚毫米级,平均位姿注册均方根误差(RMSE)为0.9毫米,尺度估计值与真实值偏差小于2%。结论:我们构建了一个模块化感知流水线,将语义分割与实时单目SLAM相结合,为自然腔道手术提供了一种有前景的场景理解方案,可促进手术自动化或辅助医生导航。

原文摘要 · Abstract (English)

Purpose: Natural orifice surgeries minimize the need for incisions and reduce the recovery time compared to open surgery; however, they require a higher level of expertise due to visualization and orientation challenges. We propose a perception pipeline for these surgeries that allows semantic scene understanding. Methods: We bring learning-based segmentation, depth estimation, and 3D reconstruction modules together to create real-time segmented maps of the surgical scenes. Additionally, we use registration with robot poses to solve the scale ambiguity of mapping from monocular images, and allow the use of semantically informed real-time reconstructions in robotic surgeries. Results: We achieve sub-milimeter reconstruction accuracy based on average one-sided Chamfer distances, average pose registration RMSE of 0.9 mm, and an estimated scale within 2% of ground truth. Conclusion: We present a modular perception pipeline, integrating semantic segmentation with real-time monocular SLAM for natural orifice surgeries. This pipeline offers a promising solution for scene understanding that can facilitate automation or surgeon guidance.

内窥镜语义建图机器人手术单目SLAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。