用视觉运动扩散模型让四指机械手学会复杂抓握操作。
Learning Dexterous In-Hand Manipulation with Multifingered Hands via Visuomotor Diffusion
- 通过增强现实界面采集高质量专家示范,实现精准远程操控。
- 采用聚类与异常检测算法过滤低质示范,提升策略学习效果。
- 在真实场景验证,支持如单手拧瓶盖等精细操作,适合机器人灵巧操作研究者。
我们提出一种基于视觉运动扩散策略的多指机械手灵巧操作学习框架。系统利用四指Allegro机械手,通过快速响应的远程操控设置完成复杂操作,如单手拧开瓶盖。采用增强现实(AR)界面追踪手部动作,结合逆运动学与运动重定向技术实现精确控制;AR头显提供实时可视化,手势控制简化操作流程。为提升策略学习效果,引入基于HDBSCAN聚类与全局-局部异常评分(GLOSH)的示范异常值剔除方法,有效过滤低质量示范数据。我们在真实世界环境中进行了广泛评估,并将所有实验视频公开于项目网站:https://dex-manip.github.io/
原文摘要 · Abstract (English)
We present a framework for learning dexterous in-hand manipulation with multifingered hands using visuomotor diffusion policies. Our system enables complex in-hand manipulation tasks, such as unscrewing a bottle lid with one hand, by leveraging a fast and responsive teleoperation setup for the four-fingered Allegro Hand. We collect high-quality expert demonstrations using an augmented reality (AR) interface that tracks hand movements and applies inverse kinematics and motion retargeting for precise control. The AR headset provides real-time visualization, while gesture controls streamline teleoperation. To enhance policy learning, we introduce a novel demonstration outlier removal approach based on HDBSCAN clustering and the Global-Local Outlier Score from Hierarchies (GLOSH) algorithm, effectively filtering out low-quality demonstrations that could degrade performance. We evaluate our approach extensively in real-world settings and provide all experimental videos on the project website: https://dex-manip.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。