arXiv:2601.01651cs.RO2026-01被引 3

从一段人拍摄的视频中,教会双臂机械手完成复杂操作。

DemoBot: Efficient Learning of Bimanual Manipulation with Dexterous Hands From Third-Person Human Videos

  • 从单个视频提取双手与物体的运动轨迹作为先验
  • 通过强化学习优化轨迹,实现长时间复杂操作
  • 适合需要快速学习人类动作的机器人研发者

本文提出DemoBot,一种从单段未标注的RGB-D视频中学习双臂多指机器人复杂操作技能的框架。该方法从原始视频中提取双手和物体的结构化运动轨迹,作为新型强化学习(RL)管道的运动先验,通过接触丰富的交互过程进行优化,避免从零学习。为解决长时序操作学习难题,引入三项关键技术:(1) 基于时间片段的强化学习,确保当前状态与示范对齐;(2) 成功触发重置策略,平衡已掌握技能的巩固与后续阶段探索;(3) 事件驱动奖励课程与自适应阈值,引导高精度操作学习。该框架成功实现了同步与异步双臂装配任务,提供了一种可扩展的直接从人类视频获取技能的方法。

原文摘要 · Abstract (English)

This work presents DemoBot, a learning framework that enables a dual-arm, multi-finger robotic system to acquire complex manipulation skills from a single unannotated RGB-D video demonstration. The method extracts structured motion trajectories of both hands and objects from raw video data. These trajectories serve as motion priors for a novel reinforcement learning (RL) pipeline that learns to refine them through contact-rich interactions, thereby eliminating the need to learn from scratch. To address the challenge of learning long-horizon manipulation skills, we introduce: (1) Temporal-segment based RL to enforce temporal alignment of the current state with demonstrations; (2) Success-Gated Reset strategy to balance the refinement of readily acquired skills and the exploration of subsequent task stages; and (3) Event-Driven Reward curriculum with adaptive thresholding to guide the RL learning of high-precision manipulation. The novel video processing and RL framework successfully achieved long-horizon synchronous and asynchronous bimanual assembly tasks, offering a scalable approach for direct skill acquisition from human videos.

双臂操作视频学习强化学习机器人技能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。