arXiv:2608.19182cs.ROcs.AI2026-08

用预训练+后训练提升机器人灵巧操作能力,实现从仿真到现实的零样本迁移。

ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning

论文配图:ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning
图 1 · 摘自论文原文
  • 先预训练再微调,利用已有技能加速新任务学习
  • 在23和29自由度机器人上实现人类级速度的长时序操作
  • 适合需要高效灵巧操作的机器人研发与工业应用

我们提出ADEPT,一种大规模强化学习框架,用于在高自由度机器人上学习可实现从仿真到现实迁移的灵巧操作,直接从原始视觉触觉感知中解决长时序任务。ADEPT先在通用物体重置任务上预训练灵巧策略,再以该预训练行为作为先验,对下游策略进行后训练。该方法使多指机器人能够学习原本难以从零开始发现的新行为,并避免为每个新任务重复学习相同技能。预训练策略能零样本完成下游任务的重置阶段,但朴素的强化学习微调会迅速破坏此能力。为此,我们提出稳定后训练方案,结合行为克隆蒸馏、评价网络热身和保守的在线更新。为安全利用完整运动学灵活性,引入关节空间几何织物作为策略与机器人间的中介。我们将后训练教师模型蒸馏为感知学生模型,在两种机器人上实现零样本仿真到现实迁移:23自由度Kuka-Allegro(双RGB相机)和29自由度Flexiv-Sharpa(双RGB相机与五个多视图触觉传感器),可在复杂初始状态下完成长时序任务,达到人类级操作速度。

原文摘要 · Abstract (English)

We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments that can solve long-horizon tasks directly from raw visuo-tactile perception. ADEPT pretrains a dexterous policy on a generic object reposing task, then post-trains downstream policies with this pretrained behavior as a prior. ADEPT enables learning new behaviors that are otherwise difficult to discover from scratch on multi-fingered robots and avoids learning the same set of skills over again for every new downstream task. The pretrained policy zero-shots the reposing phase of downstream tasks, but naïve RL fine-tuning rapidly degrades this capability during transfer. We address this with a stable post-training recipe combining behavior-cloning distillation, critic warm-up, and conservative on-policy updates. To safely exploit the full kinematic dexterity, we introduce a joint-space Geometric Fabric that mediates between the RL policy and the robot. We distill post-trained teachers into perceptive students that zero-shot sim-to-real transfer on two embodiments: a 23 DoF Kuka-Allegro with two RGB cameras, and a 29 DoF Flexiv-Sharpa with two RGB cameras and five vision-based tactile sensors, and can solve long-horizon tasks from challenging initial states with dexterity at human-level speed.

灵巧操作强化学习仿真到现实多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。