arXiv:2607.20033cs.RO2026-07被引 1

机器人仅看一段人类示范视频,30秒内学会新技能且不丢旧技能。

HOST:Robots Acquire Manipulation Skills in Seconds from a Single Human Video

论文配图:HOST:Robots Acquire Manipulation Skills in Seconds from a Single Human Video
图 1 · 摘自论文原文
  • 通过自校准预测级联,将人示范动作转为机器人可执行的未来观测
  • 单视频学习平均29秒完成技能获取,成功率62%,比零样本提升45%
  • 只需1次演示,效率远超需50次训练的模型,适合快速部署新任务

机器人快速且无损地学习新技能至关重要。现有方法依赖耗时耗力的训练循环,且会遗忘已有技能。本文提出HOST(Human-to-robot One-Shot Skill AcquisiTion)框架,使机器人仅凭一段人类示范视频,在平均29秒内完成新技能学习,并保留先前掌握的技能。HOST通过级联自校准预测:先估计任务进展,再将后续进展映射为机器人的未来观测,最后从中推导动作。该过程基于视频与机器人轨迹对齐的共享任务进展流形进行目标重构。实验表明,HOST在单视频条件下实现62%平均成功率,较零样本基线提升45%,且优于在50次机器人演示上微调的模型,但仅需1/50的演示数据,技能获取速度提升507倍。更多细节见项目网站。

原文摘要 · Abstract (English)

The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-time loop that is costly and slow, while eroding skills already mastered. In this paper, we introduce HOST (Human-to-robot One-Shot Skill AcquisiTion), a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills. HOST resolves skill acquisition through a cascade of self-grounded prediction. It first estimates the robot's progress within the demonstrated task, then translates the upcoming progression into the robot's own future observations, and finally derives actions from these predicted observations. This cascade is trained on targets coupled to the video demonstration, obtained by mapping the robot trajectory and the video demonstration onto a shared task progress manifold, then redefining each target to align with the future progression of the video. HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot's embodiment. HOST acquires novel skills at inference time from a single human video in an average of 29 seconds and achieves a 62% average success rate. It exceeds the zero-shot baseline by 45% while retaining previously mastered skills. HOST even exceeds the baseline fine-tuned on 50 robot demonstrations per task while requiring 50 times fewer demonstrations and acquiring each skill 507 times faster. Additional information about HOST is available on the project website.

机器人技能学习单次示教模仿学习快速适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。