打造可穿戴系统,采集手部操作时的视觉触觉动作数据。
DexViTac: Collecting Human Visuo-Tactile-Kinematic Demonstrations for Contact-Rich Dexterous Manipulation
- 自研可穿戴设备同步记录视觉、触觉、手部姿态与末端位姿。
- 每小时采集超248次示范,真实场景下成功率超85%。
- 适合研究灵巧操作与多模态机器人学习的研究者。
大规模高质量多模态示范数据对接触丰富的灵巧操作机器人学习至关重要。尽管以人类为中心的数据采集系统降低了规模化门槛,但在物理交互过程中难以捕捉触觉信息。为此,我们提出DexViTac,一种专为接触丰富的灵巧操作设计的便携式、以人为本的数据采集系统。该系统可在非结构化、真实环境中高保真地获取第一人称视觉、高密度触觉感知、末端位姿及手部运动学数据。基于此硬件,我们提出一种基于运动学的触觉表征学习算法,有效解决触觉信号中的语义模糊问题。利用DexViTac的高效性,我们构建了一个包含超过2,400个视觉-触觉-运动学示范的多模态数据集。实验表明,DexViTac每小时可完成超过248次示范采集,且在复杂视觉遮挡下仍保持鲁棒性。真实场景部署验证了基于该数据集和学习策略训练的策略,在四个挑战性任务中平均成功率超过85%,显著优于基线方法,验证了该系统在接触丰富灵巧操作学习中的显著提升作用。
原文摘要 · Abstract (English)
Large-scale, high-quality multimodal demonstrations are essential for robot learning of contact-rich dexterous manipulation. While human-centric data collection systems lower the barrier to scaling, they struggle to capture the tactile information during physical interactions. Motivated by this, we present DexViTac, a portable, human-centric data collection system tailored for contact-rich dexterous manipulation. The system enables the high-fidelity acquisition of first-person vision, high-density tactile sensing, end-effector poses, and hand kinematics within unstructured, in-the-wild environments. Building upon this hardware, we propose a kinematics-grounded tactile representation learning algorithm that effectively resolves semantic ambiguities within tactile signals. Leveraging the efficiency of DexViTac, we construct a multimodal dataset comprising over 2,400 visuo-tactile-kinematic demonstrations. Experiments demonstrate that DexViTac achieves a collection efficiency exceeding 248 demonstrations per hour and remains robust against complex visual occlusions. Real-world deployment confirms that policies trained with the proposed dataset and learning strategy achieve an average success rate exceeding 85% across four challenging tasks. This performance significantly outperforms baseline methods, thereby validating the substantial improvement the system provides for learning contact-rich dexterous manipulation. Project page: https://xitong-c.github.io/DexViTac/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。