用单帧投影实现无同步实时腹腔镜深度感知,精度达3.7毫米。
AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance

- 基于LED二值掩码与潜空间U-Net,无需投影同步即可单帧重建深度。
- 在722对样本上实现3.70毫米平均绝对误差,优于现有单目模型。
- 适合需实时、紧凑深度感知的微创手术机器人系统部署。
准确的术中深度感知对自主和半自主机器人腹腔镜手术至关重要。传统条纹投影轮廓术虽可达毫米级精度,但常需多帧采集、数字微镜设备投影及投影-相机同步,难以集成于紧凑型腹腔镜系统。本研究提出一种免同步、单帧深度感知平台,采用被动LED照射二值掩码与自定义U-Net深度头结合的VQ-VAE先验。将小型投影模块耦合至双通道腹腔镜的一通道,另一通道拍摄条纹照明目标。使用Zivid 3D相机获取722对伪影图像的参考深度图,并将其重投影至SSLE图像坐标系用于监督训练与评估。VQ-VAE将输入编码为离散潜变量,潜空间U-Net直接预测深度,无需独立掩码预测分支。在固定训练/验证/测试划分下,模型达到3.70毫米平均绝对误差(MAE)、0.0326绝对相对误差(AbsRel)、delta=1.1阈值准确率96.2%及delta=1.1²准确率97.0%。该模型在MAE、AbsRel和阈值准确率上均优于双分支U-Net MaskNet+DepthNet基线与现成单目深度模型。在NVIDIA A100 GPU上实现26.0赫兹运行速度,连续处理301帧。结论表明,基于LED二值图案与潜空间深度重建的方案可实现免同步、视频速率的内窥镜深度估计。结果展示无需显式分割阶段的Zivid参考伪影重建效果,同时强调数据集规模与SSLE-Zivid标定精度的重要性。
原文摘要 · Abstract (English)
Significance. Accurate intraoperative depth perception is important for autonomous and semi-autonomous robotic laparoscopic surgery. Conventional fringe projection profilometry can achieve millimeter-scale accuracy but often requires multi-shot acquisition, digital-micromirror-device projection, and projector-camera synchronization, complicating integration into compact laparoscopic systems. Aim. To develop a synchronization-free, single-shot depth-sensing platform using a passive LED-illuminated binary mask and a VQ-VAE prior with a custom U-Net depth head. Approach. A compact projection module was coupled to one channel of a dual-channel laparoscope, while the second channel imaged the fringe-illuminated target. A Zivid 3D camera acquired reference depth for 722 paired phantom images. Zivid depth maps were reprojected into the SSLE image frame for supervised training and evaluation. The VQ-VAE encoded each input into a discrete latent representation, and a latent-space U-Net predicted depth without a separate mask-prediction branch. Results. Using a fixed train/validation/test split, the proposed model achieved an MAE of 3.70 mm, AbsRel of 0.0326, delta=1.1 accuracy of 0.962, and delta=1.1^2 accuracy of 0.970. It achieved lower MAE than the dual U-Net MaskNet + DepthNet baseline and outperformed off-the-shelf monocular depth models in MAE, AbsRel, and threshold accuracy. The pipeline operated at 26.0 Hz over 301 consecutive frames on an NVIDIA A100 GPU. Conclusions. The LED-illuminated binary-pattern platform with latent-space depth reconstruction enables synchronization-free, video-rate endoscopic depth estimation. Results demonstrate Zivid-referenced phantom reconstruction without an explicit segmentation stage, while emphasizing the importance of dataset size and SSLE-Zivid calibration accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。