arXiv:2609.07498cs.ROcs.CV2026-09

将复杂手部动作精准转为机械臂抓取,提升机器人操作精度。

CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

论文配图:CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
图 1 · 摘自论文原文
  • 分两阶段生成抓取动作:先预测关键帧,再补全连续轨迹。
  • 数据集含6189次演示,覆盖1254种物体,空间动作更复杂。
  • 适合做复杂抓取任务的机器人学习研究者使用。

将人类手部示范转移到机器人夹爪已成为一种低成本的机器人学习方案。然而,现有方法多局限于简单平面任务,难以处理涉及旋转或翻转等复杂空间运动,而这类动作对机器人操作至关重要。为此,我们采用基于细粒度手部姿态的隐式数据驱动方法,设计了一套可扩展的数据采集流程,通过手持夹爪实现无缝动作模仿,并制定严格协议以保证运动复杂性。最终构建了一个大规模配对数据集,包含6,189个回合、1,254种不同物体,其空间复杂度显著高于现有基准。但学习此类复杂映射仍具挑战。我们发现直接端到端生成完整夹爪位姿序列效果不佳,因微小轨迹偏差在复杂动力学下会快速累积。为此,提出两阶段框架:第一阶段预测稀疏的夹爪关键帧(起始与终止),简化映射目标;第二阶段在关键帧约束下生成完整的连续动作序列。此外,为缓解累积漂移,保持夹爪方向学习的同时,基于抓取启发式和运动学一致性后优化其位置。仿真与真实机器人实验均表明,该框架能稳定精确地实现复杂空间操作的手-夹爪转移,显著优于传统基线方法。

原文摘要 · Abstract (English)

Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.

机器人操作动作迁移数据集抓取生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。