arXiv:2505.20904cs.CV2025-05被引 1

融合Transformer与Mamba,提升透明反射物深度补全效果

HTMNet: A Hybrid Network with Transformer-Mamba Bottleneck Multimodal Fusion for Transparent and Reflective Objects Depth Completion

  • 用CNN-Transformer双分支编码,Mamba做瓶颈融合
  • 在多个数据集上达到当前最好性能
  • 适合需要精准深度感知的机器人视觉任务

透明与反射物体对深度传感器构成重大挑战,导致深度信息不完整,影响下游机器人感知与操作。为此,我们提出HTMNet,一种融合Transformer、CNN与Mamba架构的新型混合模型。编码器采用双分支CNN-Transformer结构,瓶颈融合模块使用Transformer-Mamba架构,解码器基于多尺度融合模块。我们引入一种基于自注意力与状态空间模型的新型多模态融合机制,首次将Mamba架构应用于透明物体深度补全,展现出巨大潜力。同时,设计创新的多尺度融合模块,结合通道注意力、空间注意力与多尺度特征提取,通过下融合策略有效整合多尺度特征。在多个公开数据集上的大量评估表明,该模型达到当前最优(SOTA)性能,验证了方法的有效性。

原文摘要 · Abstract (English)

Transparent and reflective objects pose significant challenges for depth sensors, resulting in incomplete depth information that adversely affects downstream robotic perception and manipulation tasks. To address this issue, we propose HTMNet, a novel hybrid model integrating Transformer, CNN, and Mamba architectures. The encoder is based on a dual-branch CNN-Transformer framework, the bottleneck fusion module adopts a Transformer-Mamba architecture, and the decoder is built upon a multi-scale fusion module. We introduce a novel multimodal fusion module grounded in self-attention mechanisms and state space models, marking the first application of the Mamba architecture in the field of transparent object depth completion and revealing its promising potential. Additionally, we design an innovative multi-scale fusion module for the decoder that combines channel attention, spatial attention, and multi-scale feature extraction techniques to effectively integrate multi-scale features through a down-fusion strategy. Extensive evaluations on multiple public datasets demonstrate that our model achieves state-of-the-art(SOTA) performance, validating the effectiveness of our approach.

深度补全多模态融合Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。