arXiv:2509.16532cs.ROcs.AI2025-09

用2D图像生成伪3D特征,让机器人在不依赖真实3D数据下实现高效抓取

No Need for Real 3D: Fusing 2D Vision with Pseudo 3D Representations for Robotic Manipulation Learning

  • 将单目图像转为保留几何结构的伪点云,与2D特征融合
  • 在多个任务上达到与真实3D方法相当的性能
  • 适合无3D传感器或追求低成本部署的机器人场景

近年来,基于视觉的机器人操作受到广泛关注并取得显著进展。当前主流方法分为基于2D图像和基于3D点云的策略学习,后者在性能和泛化能力上均更优,凸显了3D信息的价值。然而,3D点云获取成本高,限制了其可扩展性与实际部署。为此,我们提出NoReal3D框架:引入可学习的3DStructureFormer模块,将单目图像转化为具有几何意义的伪点云特征,并与2D编码器输出融合。特别地,生成的伪点云保持几何与拓扑结构,因此设计了专门的伪点云编码器以保留这些特性。我们还研究了不同特征融合策略的有效性。该框架在无需真实3D点云的情况下,显著提升了机器人对3D空间结构的理解。大量实验表明,本方法在多种任务中性能可媲美基于真实3D点云的方法。

原文摘要 · Abstract (English)

Recently,vision-based robotic manipulation has garnered significant attention and witnessed substantial advancements. 2D image-based and 3D point cloud-based policy learning represent two predominant paradigms in the field, with recent studies showing that the latter consistently outperforms the former in terms of both policy performance and generalization, thereby underscoring the value and significance of 3D information. However, 3D point cloud-based approaches face the significant challenge of high data acquisition costs, limiting their scalability and real-world deployment. To address this issue, we propose a novel framework NoReal3D: which introduces the 3DStructureFormer, a learnable 3D perception module capable of transforming monocular images into geometrically meaningful pseudo-point cloud features, effectively fused with the 2D encoder output features. Specially, the generated pseudo-point clouds retain geometric and topological structures so we design a pseudo-point cloud encoder to preserve these properties, making it well-suited for our framework. We also investigate the effectiveness of different feature fusion strategies.Our framework enhances the robot's understanding of 3D spatial structures while completely eliminating the substantial costs associated with 3D point cloud acquisition.Extensive experiments across various tasks validate that our framework can achieve performance comparable to 3D point cloud-based methods, without the actual point cloud data.

机器人操作伪3D视觉感知少数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。