arXiv:2603.13869cs.ROcs.AI2026-03

用点云重建提升抓取透明物体的精度和鲁棒性

TransDex: Pre-training Visuo-Tactile Policy with Point Cloud Reconstruction for Dexterous Manipulation of Transparent Objects

  • 通过Transformer自监督学习重建手部交互点云的3D结构
  • 在加噪和大范围遮挡下仍能准确恢复物体三维形态
  • 适合需要精细操作透明物体的机器人系统

灵巧操作复杂任务时,透明物体常因自遮挡、深度噪声和深度信息丢失导致失败。本文提出TransDex,一种基于点云重建预训练的3D视觉-触觉融合控制策略。首先设计了一种基于Transformer的自监督点云重建方法,即使在随机噪声和大范围遮挡下,也能从灵巧手交互点云中精确恢复物体3D结构。在此基础上,TransDex采用细粒度分层感知编码与多轮注意力机制,自适应融合机械臂与灵巧手特征,实现差异化运动预测。真实机器人系统上的透明物体操作实验表明,TransDex优于现有基线方法。进一步分析验证了其泛化能力及各组件的有效性。

原文摘要 · Abstract (English)

Dexterous manipulation enables complex tasks but suffers from self-occlusion, severe depth noise, and depth information loss when manipulating transparent objects. To solve this problem, this paper proposes TransDex, a 3D visuo-tactile fusion motor policy based on point cloud reconstruction pre-training. Specifically, we first propose a self-supervised point cloud reconstruction pre-training approach based on Transformer. This method accurately recovers the 3D structure of objects from interactive point clouds of dexterous hands, even when random noise and large-scale masking are added. Building on this, TransDex is constructed in which perceptual encoding adopts a fine-grained hierarchical scheme and multi-round attention mechanisms adaptively fuse features of the robotic arm and dexterous hand to enable differentiated motion prediction. Results from transparent object manipulation experiments conducted on a real robotic system demonstrate that TransDex outperforms existing baseline methods. Further analysis validates the generalization capabilities of TransDex and the effectiveness of its individual components.

灵巧操作点云重建多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。