arXiv:2504.17401cs.CVcs.AI2025-04被引 1

提出 StereoMamba 模型,实现手术中实时高精度立体视差估计。

StereoMamba: Real-time and Robust Intraoperative Stereo Disparity Estimation via Long-range Spatial Dependencies

  • 设计 FE-Mamba 模块捕捉图像内外的长距离空间依赖。
  • 在 SCARED 基准上达到 2.64px EPE 与 21.28 FPS 实时性能。
  • 零样本泛化能力强,适用于真实手术场景数据集。

立体视差估计对于获取机器人辅助微创手术(RAMIS)中的深度信息至关重要。尽管现有深度学习方法已取得显著进展,但在准确性、鲁棒性与推理速度之间仍难以平衡。为此,我们提出 StereoMamba 架构,专为 RAMIS 中的立体视差估计设计。其基于新型特征提取模块 FE-Mamba,增强单张及双目图像间的长程空间依赖。为进一步融合多尺度特征,引入多维特征融合(MFF)模块。在体外 SCARED 基准上的实验表明,StereoMamba 在 EPE 上达 2.64 px,深度 MAE 为 2.55 mm,Bad2 为 41.49%,Bad3 为 26.99%,同时保持 21.28 FPS 推理速度(1280×1024 高分辨率图像对),实现准确、鲁棒与高效的最佳平衡。通过将合成右图(由左图经视差图投影生成)与真实右图对比,StereoMamba 在体内 RIS2017 与 StereoMIS 数据集上取得最优平均 SSIM(0.8970)与 PSNR(16.0761),展现出优异的零样本泛化能力。

原文摘要 · Abstract (English)

Stereo disparity estimation is crucial for obtaining depth information in robot-assisted minimally invasive surgery (RAMIS). While current deep learning methods have made significant advancements, challenges remain in achieving an optimal balance between accuracy, robustness, and inference speed. To address these challenges, we propose the StereoMamba architecture, which is specifically designed for stereo disparity estimation in RAMIS. Our approach is based on a novel Feature Extraction Mamba (FE-Mamba) module, which enhances long-range spatial dependencies both within and across stereo images. To effectively integrate multi-scale features from FE-Mamba, we then introduce a novel Multidimensional Feature Fusion (MFF) module. Experiments against the state-of-the-art on the ex-vivo SCARED benchmark demonstrate that StereoMamba achieves superior performance on EPE of 2.64 px and depth MAE of 2.55 mm, the second-best performance on Bad2 of 41.49% and Bad3 of 26.99%, while maintaining an inference speed of 21.28 FPS for a pair of high-resolution images (1280*1024), striking the optimum balance between accuracy, robustness, and efficiency. Furthermore, by comparing synthesized right images, generated from warping left images using the generated disparity maps, with the actual right image, StereoMamba achieves the best average SSIM (0.8970) and PSNR (16.0761), exhibiting strong zero-shot generalization on the in-vivo RIS2017 and StereoMIS datasets.

立体视觉手术导航实时估计Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。