arXiv:2511.07081cs.RO2025-11被引 2

融合Transformer、CNN与Mamba,提升透明反光物体的深度补全精度。

HDCNet: A Hybrid Depth Completion Network for Grasping Transparent and Reflective Objects

  • 双分支结构提取多模态特征,浅层轻量融合低级信息。
  • 瓶颈处采用Transformer-Mamba混合模块,增强语义与上下文理解。
  • 实测抓取成功率最高提升60%,适合机器人感知任务。

透明与反光物体的深度感知长期是机器人操作中的关键挑战。传统深度传感器在这些表面常无法提供可靠测量,制约了机器人在感知与抓取任务中的表现。为此,我们提出一种新型深度补全网络HDCNet,融合Transformer、CNN与Mamba架构的优势。编码器采用双分支Transformer-CNN结构,提取模态特异性特征;浅层引入轻量级多模态融合模块,有效整合低层特征;在网络瓶颈处设计Transformer-Mamba混合融合模块,实现高层语义与全局上下文信息的深度融合,显著提升深度补全的精度与鲁棒性。在多个公开数据集上的大量评估表明,HDCNet在深度补全任务中达到当前最优(SOTA)性能。此外,机器人抓取实验显示,该方法显著提升对透明与反光物体的抓取成功率,最高可达60%的提升。

原文摘要 · Abstract (English)

Depth perception of transparent and reflective objects has long been a critical challenge in robotic manipulation.Conventional depth sensors often fail to provide reliable measurements on such surfaces, limiting the performance of robots in perception and grasping tasks. To address this issue, we propose a novel depth completion network,HDCNet,which integrates the complementary strengths of Transformer,CNN and Mamba architectures.Specifically,the encoder is designed as a dual-branch Transformer-CNN framework to extract modality-specific features. At the shallow layers of the encoder, we introduce a lightweight multimodal fusion module to effectively integrate low-level features. At the network bottleneck,a Transformer-Mamba hybrid fusion module is developed to achieve deep integration of high-level semantic and global contextual information, significantly enhancing depth completion accuracy and robustness. Extensive evaluations on multiple public datasets demonstrate that HDCNet achieves state-of-the-art(SOTA) performance in depth completion tasks.Furthermore,robotic grasping experiments show that HDCNet substantially improves grasp success rates for transparent and reflective objects,achieving up to a 60% increase.

深度补全机器人抓取多模态融合Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。