首个融合视觉触觉本体感觉的大规模手握物体姿态数据集
VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception
- 构建多模态数据集,整合视觉、触觉与本体感知信号
- 包含200万仿真与10万真实数据,支持高精度手部操作建模
- 适合研究机器人抓取、多模态融合及仿真到真实的迁移
本文针对手握物体姿态估计缺乏大规模数据集的问题,提出VinT-6D——首个融合视觉、触觉与本体感知的多模态大规模数据集,以提升机器人在感知-规划-控制框架下的手部操作能力。该数据集包含200万条仿真数据(VinT-Sim)和10万条真实世界数据(VinT-Real),通过MuJoCo与Blender仿真平台以及自研真实实验平台采集。数据专为机械手设计,具备全手触觉感知与高质量对齐特性。据我们所知,VinT-Real是当前真实环境下规模最大的同类数据集,有助于弥合仿真与现实之间的差距。基于此数据集,我们提出了一个融合多模态信息的基准方法,在性能上实现显著提升。项目主页:https://VinT-6D.github.io/
原文摘要 · Abstract (English)
This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the ``Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision, touch, and proprioception, to enhance robotic manipulation. VinT-6D comprises 2 million VinT-Sim and 0.1 million VinT-Real splits, collected via simulations in MuJoCo and Blender and a custom-designed real-world platform. This dataset is tailored for robotic hands, offering models with whole-hand tactile perception and high-quality, well-aligned data. To the best of our knowledge, the VinT-Real is the largest considering the collection difficulties in the real-world environment so that it can bridge the gap of simulation to real compared to the previous works. Built upon VinT-6D, we present a benchmark method that shows significant improvements in performance by fusing multi-modal information. The project is available at https://VinT-6D.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。