用单目视频重建可控制的手术器械数字孪生,提升仿真与训练效果。
Instrument-Splatting++: Towards Controllable Surgical Instrument Digital Twin Using Gaussian Splatting
- 基于3D高斯点云,注入CAD先验实现分部件语义建模。
- 从无标注内窥镜视频中恢复6自由度姿态和关节角度,精度优于现有方法。
- 适合机器人手术仿真、合成数据生成与新视角泛化任务使用。
高质量且可控的手术器械数字孪生对机器人辅助手术中的真实世界到模拟世界(Real2Sim)至关重要,可支持逼真的仿真、合成数据生成及感知学习。本文提出Instrument-Splatting++,一种单目三维高斯点云(3DGS)框架,将手术器械重建为高保真、完全可控的高斯资产。其流程始于分部件几何预训练,将CAD先验注入高斯基元,赋予表示以部件感知的语义渲染能力。在此基础上,提出语义感知姿态估计与跟踪(SAPET)方法,从无标注内窥镜视频中恢复每帧6-DoF姿态与关节角,其中纯合成语义训练的钳端网络提供鲁棒监督,松散正则化抑制奇异运动。最后引入鲁棒纹理学习(RTL),通过姿态精修与外观优化交替进行,缓解纹理学习过程中的姿态噪声。该框架可从无标注视频中完成姿态估计并学习真实纹理。我们在EndoVis17/18、SAR-RARP及自建数据集上验证,相比现有最优基线,光照质量更优,几何精度更高。进一步在下游关键点检测任务中,利用可控器械高斯模型生成未见姿态数据,显著提升性能。
原文摘要 · Abstract (English)
High-quality and controllable digital twins of surgical instruments are critical for Real2Sim in robot-assisted surgery, as they enable realistic simulation, synthetic data generation, and perception learning under novel poses. We present Instrument-Splatting++, a monocular 3D Gaussian Splatting (3DGS) framework that reconstructs surgical instruments as a fully controllable Gaussian asset with high fidelity. Our pipeline starts with part-wise geometry pretraining that injects CAD priors into Gaussian primitives and equips the representation with part-aware semantic rendering. Built on the pretrained model, we propose a semantics-aware pose estimation and tracking (SAPET) method to recover per-frame 6-DoF pose and joint angles from unposed endoscopic videos, where a gripper-tip network trained purely from synthetic semantics provides robust supervision and a loose regularization suppresses singular articulations. Finally, we introduce Robust Texture Learning (RTL), which alternates pose refinement and robust appearance optimization, mitigating pose noise during texture learning. The proposed framework can perform pose estimation and learn realistic texture from unposed videos. We validate our method on sequences extracted from EndoVis17/18, SAR-RARP, and an in-house dataset, showing superior photometric quality and improved geometric accuracy over state-of-the-art baselines. We further demonstrate a downstream keypoint detection task where unseen-pose data augmentation from our controllable instrument Gaussian improves performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。