arXiv:2503.17106cs.CVcs.RO2025-03

针对透明反光物体深度缺失问题,提出融合3D几何结构的补全方法。

GAA-TSO: Geometry-Aware Assisted Depth Completion for Transparent and Specular Objects

  • 引入点云构建3D分支,提取场景级几何特征。
  • 通过门控融合模块将3D特征有效传递至2D图像分支。
  • 在多个数据集上提升深度补全与机器人抓取性能。

透明和镜面物体在日常生活、工厂和实验室中常见,但其独特的光学特性导致深度信息常不完整且不准确,严重影响下游机器人任务。现有方法依赖RGB信息辅助深度预测,但因这些物体纹理差,易生成无结构的深度图,且2D方法难以挖掘深度通道中的3D结构,造成深度歧义。为此,我们提出一种几何感知辅助的深度补全方法,重点利用场景的3D结构线索。除从RGB-D输入提取2D特征外,还将输入深度回投影为点云,建立3D分支以提取分层的场景级3D结构特征。设计多个门控跨模态融合模块,有效将多层级3D几何特征传播至图像分支。此外,提出自适应相关性聚合策略,合理分配3D特征到对应2D特征。在ClearGrasp、OOD、TransCG和STD数据集上的大量实验表明,本方法优于现有最先进方法。进一步验证其显著提升下游机器人抓取任务性能。

原文摘要 · Abstract (English)

Transparent and specular objects are frequently encountered in daily life, factories, and laboratories. However, due to the unique optical properties, the depth information on these objects is usually incomplete and inaccurate, which poses significant challenges for downstream robotics tasks. Therefore, it is crucial to accurately restore the depth information of transparent and specular objects. Previous depth completion methods for these objects usually use RGB information as an additional channel of the depth image to perform depth prediction. Due to the poor-texture characteristics of transparent and specular objects, these methods that rely heavily on color information tend to generate structure-less depth predictions. Moreover, these 2D methods cannot effectively explore the 3D structure hidden in the depth channel, resulting in depth ambiguity. To this end, we propose a geometry-aware assisted depth completion method for transparent and specular objects, which focuses on exploring the 3D structural cues of the scene. Specifically, besides extracting 2D features from RGB-D input, we back-project the input depth to a point cloud and build the 3D branch to extract hierarchical scene-level 3D structural features. To exploit 3D geometric information, we design several gated cross-modal fusion modules to effectively propagate multi-level 3D geometric features to the image branch. In addition, we propose an adaptive correlation aggregation strategy to appropriately assign 3D features to the corresponding 2D features. Extensive experiments on ClearGrasp, OOD, TransCG, and STD datasets show that our method outperforms other state-of-the-art methods. We further demonstrate that our method significantly enhances the performance of downstream robotic grasping tasks.

深度补全3D几何机器人抓取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。