arXiv:2512.24210cs.RO2025-12被引 8

让双臂灵巧手机器人能听懂指令完成复杂操作

GR-Dexter Technical Report

  • 设计21自由度灵巧手+远程操控系统收集真实数据
  • 在多种任务中实现强性能和对未知物体的鲁棒性
  • 适合研究通用灵巧操作的机器人学者参考

视觉-语言-动作(VLA)模型已实现语言驱动的长时序机器人操作,但多数系统仅限于夹爪。将VLA策略扩展至高自由度双臂灵巧手机器人仍具挑战,主要因动作空间扩大、手物遮挡频繁以及真实数据采集成本高。本文提出GR-Dexter,一个面向双臂灵巧手机器人的完整软硬件-数据框架。该方案结合紧凑的21-DoF机械手设计、直观的双臂远程操控系统用于真实数据采集,以及融合远程操控轨迹、大规模视觉-语言数据与精心筛选的跨体态数据集的训练方法。在真实世界评估中,GR-Dexter在长时序日常操作与可泛化的拾取放置任务中均表现出色,具备对未见物体和未见指令的更强鲁棒性。我们希望GR-Dexter成为通用灵巧手机器人操作的一次实用进展。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have enabled language-conditioned, long-horizon robot manipulation, but most existing systems are limited to grippers. Scaling VLA policies to bimanual robots with high degree-of-freedom (DoF) dexterous hands remains challenging due to the expanded action space, frequent hand-object occlusions, and the cost of collecting real-robot data. We present GR-Dexter, a holistic hardware-model-data framework for VLA-based generalist manipulation on a bimanual dexterous-hand robot. Our approach combines the design of a compact 21-DoF robotic hand, an intuitive bimanual teleoperation system for real-robot data collection, and a training recipe that leverages teleoperated robot trajectories together with large-scale vision-language and carefully curated cross-embodiment datasets. Across real-world evaluations spanning long-horizon everyday manipulation and generalizable pick-and-place, GR-Dexter achieves strong in-domain performance and improved robustness to unseen objects and unseen instructions. We hope GR-Dexter serves as a practical step toward generalist dexterous-hand robotic manipulation.

灵巧操作双臂机器人视觉语言动作数据采集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。