arXiv:2509.09671cs.ROcs.CV2025-09被引 11

用强化学习统一优化抓取动作,让机器人从不完美演示中学会灵巧操作。

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration

  • 单循环优化同时完成动作重定向与跟踪,避免多阶段误差累积。
  • 在大规模演示数据上训练出低控制代价、抗噪声的机器人策略。
  • 适合需要灵巧操作且有真实演示数据的研究者或工程师。

手-物运动捕捉(MoCap)数据集提供大量接触丰富的示范,有望推动灵巧机器人操作的规模化发展。然而,示范中的误差及人手与机器人之间的形态差异限制了直接使用这些数据。现有方法采用重定向、跟踪和残差修正三阶段流程,常导致示范数据利用不足且误差逐级放大。我们提出Dexplore,一种统一的单循环优化框架,直接从原始MoCap数据中联合进行动作重定向与跟踪,学习机器人控制策略。不将示范视为真值,而是作为软性指导,从原始轨迹中推导自适应空间范围,通过强化学习使策略保持在范围内,同时最小化控制代价并完成任务。该方法保留示范意图,支持机器人特异性策略涌现,提升对噪声的鲁棒性,并可扩展至大规模示范数据集。我们将训练好的跟踪策略提炼为基于视觉、技能条件化的生成控制器,以丰富潜在表示编码多样操作技能,支持跨物体泛化与真实场景部署。总体而言,Dexplore为将不完美示范转化为有效训练信号提供了系统性解决方案。

原文摘要 · Abstract (English)

Hand-object motion-capture (MoCap) repositories offer large-scale, contact-rich demonstrations and hold promise for scaling dexterous robotic manipulation. Yet demonstration inaccuracies and embodiment gaps between human and robot hands limit the straightforward use of these data. Existing methods adopt a three-stage workflow, including retargeting, tracking, and residual correction, which often leaves demonstrations underused and compound errors across stages. We introduce Dexplore, a unified single-loop optimization that jointly performs retargeting and tracking to learn robot control policies directly from MoCap at scale. Rather than treating demonstrations as ground truth, we use them as soft guidance. From raw trajectories, we derive adaptive spatial scopes, and train with reinforcement learning to keep the policy in-scope while minimizing control effort and accomplishing the task. This unified formulation preserves demonstration intent, enables robot-specific strategies to emerge, improves robustness to noise, and scales to large demonstration corpora. We distill the scaled tracking policy into a vision-based, skill-conditioned generative controller that encodes diverse manipulation skills in a rich latent representation, supporting generalization across objects and real-world deployment. Taken together, these contributions position Dexplore as a principled bridge that transforms imperfect demonstrations into effective training signals for dexterous manipulation.

灵巧操作强化学习动作重定向生成控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。