用3D高斯点云实现单目视觉的灵巧手物体翻转,提升真实场景适应性。
ViserDex: Visual Sim-to-Real for Robust Dexterous In-hand Reorientation
- 在高斯表示空间做域随机化,生成逼真渲染数据
- 单目RGB下实现5种物体在复杂光照下的稳定翻转
- 可在消费级硬件上独立训练感知与控制模型
灵巧手内物体翻转需要精确的姿态估计以应对复杂的任务动态。虽然RGB传感能提供丰富的语义线索用于姿态跟踪,但现有方案依赖多相机系统或昂贵的光线追踪。我们提出一种单目RGB的模拟到现实框架,结合3D高斯溅射(3DGS)弥合视觉模拟到现实的差距。核心思路是在高斯表示空间进行域随机化:通过对3D高斯应用物理一致的预渲染增强,生成用于姿态估计的逼真、随机化的视觉数据。操控策略通过基于课程的强化学习与教师-学生蒸馏训练,实现复杂行为的高效学习。重要的是,感知和控制模型均可在消费级硬件上独立训练,无需大型计算集群。实验表明,使用3DGS数据训练的姿态估计算法在挑战性视觉环境下优于传统渲染数据训练的结果。我们在配备RGB相机的物理多指手系统上验证了该系统,成功实现了五种不同物体在复杂光照条件下的鲁棒翻转。结果表明,高斯溅射是实现仅依赖RGB的灵巧操作的一条实用路径。视频及附加材料详见项目网站:https://rffr.leggedrobotics.com/works/viserdex/
原文摘要 · Abstract (English)
In-hand object reorientation requires precise estimation of the object pose to handle complex task dynamics. While RGB sensing offers rich semantic cues for pose tracking, existing solutions rely on multi-camera setups or costly ray tracing. We present a sim-to-real framework for monocular RGB in-hand reorientation that integrates 3D Gaussian Splatting (3DGS) to bridge the visual sim-to-real gap. Our key insight is performing domain randomization in the Gaussian representation space: by applying physically consistent, pre-rendering augmentations to 3D Gaussians, we generate photorealistic, randomized visual data for object pose estimation. The manipulation policy is trained using curriculum-based reinforcement learning with teacher-student distillation, enabling efficient learning of complex behaviors. Importantly, both perception and control models can be trained independently on consumer-grade hardware, eliminating the need for large compute clusters. Experiments show that the pose estimator trained with 3DGS data outperforms those trained using conventional rendering data in challenging visual environments. We validate the system on a physical multi-fingered hand equipped with an RGB camera, demonstrating robust reorientation of five diverse objects even under challenging lighting conditions. Our results highlight Gaussian splatting as a practical path for RGB-only dexterous manipulation. For videos of the hardware deployments and additional supplementary materials, please refer to the project website: https://rffr.leggedrobotics.com/works/viserdex/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。