用3D视觉线索让手术机器人看清空间,无需额外摄像头。
Learning Surgical Robotic Manipulation with 3D Spatial Priors
- 直接从内窥镜图像提取3D空间特征,端到端训练。
- 在30K对立体图像上训练,实测打结和离体器官分离任务领先。
- 适合想提升手术机器人空间感知能力的研究者与工程师。
实现3D空间感知对精密手术机器人操作至关重要。现有方法或先重建场景,或依赖腕部附加摄像头增强多视角特征,但前者因多阶段处理易累积误差且难以端到端优化,后者则因干扰机械臂运动难用于临床。本文提出空间手术变压器(SST),一种端到端视觉-动作策略,通过直接挖掘内窥镜图像中的3D空间线索赋予机器人空间感知能力。首先构建了包含3万对立体内窥镜图像的大型逼真数据集Surgical3D,其具有精确3D几何信息;基于此,微调一个强大的几何变换器以从立体图像中提取鲁棒3D隐式表示。这些表示通过轻量级多层次空间特征连接器(MSFC)与机器人动作空间无缝对齐,全部在以内窥镜为中心的坐标系中完成。大量真实机器人实验表明,SST在复杂手术任务如打结和离体器官分离上达到当前最优性能,并具备强空间泛化能力,标志着向临床部署迈出重要一步。数据集与代码将公开。
原文摘要 · Abstract (English)
Achieving 3D spatial awareness is crucial for surgical robotic manipulation, where precise and delicate operations are required. Existing methods either explicitly reconstruct the surgical scene prior to manipulation, or enhance multi-view features by adding wrist-mounted cameras to supplement the default stereo endoscopes. However, both paradigms suffer from notable limitations: the former easily leads to error accumulation and prevents end-to-end optimization due to its multi-stage nature, while the latter is rarely adopted in clinical practice since wrist-mounted cameras can interfere with the motion of surgical robot arms. In this work, we introduce the Spatial Surgical Transformer (SST), an end-to-end visuomotor policy that empowers surgical robots with 3D spatial awareness by directly exploring 3D spatial cues embedded in endoscopic images. First, we build Surgical3D, a large-scale photorealistic dataset containing 30K stereo endoscopic image pairs with accurate 3D geometry, addressing the scarcity of 3D data in surgical scenes. Based on Surgical3D, we finetune a powerful geometric transformer to extract robust 3D latent representations from stereo endoscopes images. These representations are then seamlessly aligned with the robot's action space via a lightweight multi-level spatial feature connector (MSFC), all within an endoscope-centric coordinate frame. Extensive real-robot experiments demonstrate that SST achieves state-of-the-art performance and strong spatial generalization on complex surgical tasks such as knot tying and ex-vivo organ dissection, representing a significant step toward practical clinical deployment. The dataset and code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。