针对模糊组织边界,提出实时端到端内窥镜立体匹配方法
Robust Real-Time Endoscopic Stereo Matching under Fuzzy Tissue Boundaries
- 引入3D Mamba坐标注意力模块,增强位置敏感特征聚合
- 在SCARED和SERV-CT上达到42帧/秒实时速度,精度领先
- 适合微创手术机器人视觉系统开发人员使用
实时获取准确的场景深度对自动化机器人微创手术至关重要。双目内窥镜立体匹配可提供该深度信息。然而,现有立体匹配方法主要面向自然图像设计,常因组织边界模糊而失效,且难以满足高分辨率内窥镜图像的实时需求。为此,我们提出专用于内窥镜图像的实时立体匹配方法RRESM。该方法融合3D Mamba坐标注意力模块,通过位置敏感注意力图与Mamba块实现长程空间依赖建模,生成鲁棒代价体积且计算开销低。此外,引入高频视差优化模块,在小波域放大高频细节,提升组织边界附近视差预测精度。在SCARED和SERV-CT数据集上的评估表明,该方法达到当前最优匹配精度,推理速度达42 FPS。代码已开源。
原文摘要 · Abstract (English)
Real-time acquisition of accurate scene depth is essential for automated robotic minimally invasive surgery. Stereo matching with binocular endoscopy can provide this depth information. However, existing stereo matching methods, designed primarily for natural images, often struggle with endoscopic images due to fuzzy tissue boundaries and typically fail to meet real-time requirements for high-resolution endoscopic image inputs. To address these challenges, we propose \textbf{RRESM}, a real-time stereo matching method tailored for endoscopic images. Our approach integrates a 3D Mamba Coordinate Attention module that enhances cost aggregation through position-sensitive attention maps and long-range spatial dependency modeling via the Mamba block, generating a robust cost volume without substantial computational overhead. Additionally, we introduce a High-Frequency Disparity Optimization module that refines disparity predictions near tissue boundaries by amplifying high-frequency details in the wavelet domain. Evaluations on the SCARED and SERV-CT datasets demonstrate state-of-the-art matching accuracy with a real-time inference speed of 42 FPS. The code is available at https://github.com/Sonne-Ding/RRESM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。