arXiv:2502.20396cs.ROcs.AI2025-02被引 87

用仿真训练让人形机器人学会视觉抓取,效果好且可推广。

Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids

  • 通过仿真到现实的强化学习,结合接触与目标奖励机制。
  • 在未见过物体上实现高成功率,策略自适应性强。
  • 适合研究人形机器人操控或想做真实场景部署的团队。

学习通用的机器人操控策略,尤其是复杂多指人形机器人,仍是重大挑战。现有方法主要依赖大量数据采集和模仿学习,成本高、难扩展。仿真到现实的强化学习(Sim-to-Real RL)提供了替代路径,但多限于简单状态或单手任务。本文提出一种实用的仿真到现实强化学习方案,训练人形机器人完成三项高难度灵巧操作任务:抓取并移动、箱子抬起、双手交接。方法包括自动化的现实到仿真调参模块、基于接触与物体目标的泛化奖励设计、分而治之的策略蒸馏框架,以及模态特定增强的混合物体表征策略。实验表明,在未见过物体上仍能保持高成功率,策略表现出强鲁棒性和自适应性,证明视觉驱动的灵巧操控通过仿真到现实强化学习不仅可行,而且可扩展、适用于真实人形机器人任务。

原文摘要 · Abstract (English)

Learning generalizable robot manipulation policies, especially for complex multi-fingered humanoids, remains a significant challenge. Existing approaches primarily rely on extensive data collection and imitation learning, which are expensive, labor-intensive, and difficult to scale. Sim-to-real reinforcement learning (RL) offers a promising alternative, but has mostly succeeded in simpler state-based or single-hand setups. How to effectively extend this to vision-based, contact-rich bimanual manipulation tasks remains an open question. In this paper, we introduce a practical sim-to-real RL recipe that trains a humanoid robot to perform three challenging dexterous manipulation tasks: grasp-and-reach, box lift and bimanual handover. Our method features an automated real-to-sim tuning module, a generalized reward formulation based on contact and object goals, a divide-and-conquer policy distillation framework, and a hybrid object representation strategy with modality-specific augmentation. We demonstrate high success rates on unseen objects and robust, adaptive policy behaviors -- highlighting that vision-based dexterous manipulation via sim-to-real RL is not only viable, but also scalable and broadly applicable to real-world humanoid manipulation tasks.

人形机器人灵巧操作仿真到现实强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。