单张图像预测透明物体三维占据,实现真实机器人抓取。
Trans2Occ: Voxel Occupancy Estimation and Grasp for Transparent Objects from Simulation to Reality

- 从单张图像直接预测体素占据,构建几何感知表示。
- 仿真训练后无需微调,直接迁移到真实场景表现稳定。
- 基于占据信息的规则抓取策略,对透明物有效且易部署。
透明物体因折射和反射导致深度感知不可靠,仍是机器人感知的难题。现有方法依赖多视角重建或深度补全,难以规模化部署。本文提出一种基于单视图RGB输入的实用框架,直接从图像预测体素空间占据,生成支持下游抓取的几何感知表示。为支持大规模训练,构建仿真管道,在多种材质与光照条件下生成配对的RGB图像与体素占据标注。实验表明,该占据表示对领域偏移具有鲁棒性,可有效从仿真迁移至真实机器人系统,无需微调。在此基础上设计的简单规则抓取策略,实现了对透明物体的可靠抓取。仿真与真实环境中的大量实验验证了框架在3D理解与实际操作上的有效性。结果表明,单视图占据预测为机器人中透明物体感知提供了可扩展、高效的解决方案。
原文摘要 · Abstract (English)
Transparent objects remain challenging for robotic perception due to unreliable depth sensing caused by refraction and reflection. While prior approaches rely on multi-view reconstruction or depth completion, they are often difficult to scale or deploy in real-world robotic systems. In this paper, we present a practical framework for transparent object perception and manipulation based on single-view RGB input. Our approach predicts voxel-space occupancy directly from a single image, providing a geometry-aware representation that supports downstream robotic grasping. To enable large-scale training, we construct a simulation pipeline that generates paired RGB images and voxel occupancy annotations under diverse materials and lighting conditions. We demonstrate that the predicted occupancy representation is robust to domain shifts and transfers effectively from simulation to real-world robotic setups without fine-tuning. A simple rule-based grasping strategy built on top of the occupancy further achieves reliable grasp performance on transparent objects. Extensive experiments in both simulation and real-world environments show that our framework provides accurate 3D understanding and enables practical manipulation of transparent objects. These results suggest that single-view occupancy prediction offers a scalable and effective solution for transparent object perception in robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。