用隐空间先验引导机械手操作,让真实世界学习更稳定高效
LAMP: Latent Motion Prior-Guided Real-World Learning for Dexterous Hand Manipulation

- 构建隐动作先验模块,将历史动作映射为紧凑的潜在指令
- 从少量演示数据出发,经在线强化学习后成功率提升至98.75%
- 适合需要高精度操控的真实机器人任务,尤其擅长减少接触丢失
真实世界中灵巧手操作的学习仍易失败,因高维动作空间放大模仿误差,且强化学习探索常导致接触中断。现有方法虽结合模仿学习(IL)与在线强化学习(RL),但直接在原始动作空间探索效率低、对物理硬件风险高。本文提出潜动先验模块(\ extit{prior}{}),将近期动作历史映射到一个紧凑、条件依赖于历史的潜空间,并解码为可执行的高维手部目标。基于此,提出三阶段真实世界灵巧操作框架:首先从示范数据预训练潜先验,其次训练视觉-运动策略预测原始臂部命令和潜手动作偏移,最后在相同潜空间内进行残差强化学习。该共享可解码接口使探索仅在已演示的稳定接触动作附近做局部修正,而非独立扰动每个手指关节。在四个真实机器人灵巧操作任务上评估,相比原始、线性及离散动作接口,本文方法从少量特定任务示范开始,模仿学习成功率平均达56.25%,经在线强化学习后提升至98.75%;其中三项任务达到100%最终成功率,另一项为95%。
原文摘要 · Abstract (English)
Real-world learning for dexterous hands remains brittle because high-dimensional hand actions amplify imitation errors and make reinforcement-learning exploration prone to contact-breaking motion. While combining imitation learning (IL) with online reinforcement learning (RL) can reduce manual supervision, unconstrained exploration in raw hand-action spaces is sample-inefficient and risky for physical hardware. We introduce a latent motion prior module (\prior{}) that maps recent hand-action histories to a compact, history-conditioned latent prior and decodes continuous latent commands into executable high-dimensional hand targets. Built on this prior, \method{} is a three-stage real-world dexterous learning framework: it pretrains \prior{} from demonstrations, trains a visuomotor policy that predicts native arm commands and latent hand-action offsets, and improves the policy with online residual RL in the same latent hand-action space. This shared, decodable interface lets residual exploration make local corrections near demonstrated, contact-consistent hand motions rather than perturbing every finger joint independently. We evaluate \method{} on four real-robot dexterous manipulation tasks against raw, linear, and discrete hand-action interfaces. Starting from small task-specific demonstration sets, \method{} achieves a 56.25\% average IL success rate and raises it to 98.75\% after online RL, reaching 100\% final success on three tasks and 95\% on the remaining task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。