arXiv:2606.21406cs.ROcs.CV2026-06被引 2

用人类视频学动作与反馈,让机器人自动纠错提升成功率。

Robot Self-Improvement via Human-Video Dynamics Models

论文配图:Robot Self-Improvement via Human-Video Dynamics Models
图 1 · 摘自论文原文
  • 从人类视频中提取通用动作与动态模型,跨机器人形态复用。
  • 失败后自动生成并排序纠正动作,成功率从40%提升至81%。
  • 无需额外训练,适合想实现自进化能力的机器人研究者。

机器人学习的核心问题是如何从人类学习的数据类型中获取技能:被动观察、具身实践和失败经验。人类视频提供了丰富的观察数据,已有研究证明其可初始化有效策略。但更不清楚的是,这些视频能否支撑机器人的实践与纠错。本文表明,从人类视频中学习的动作、动力学和价值表征具有不依赖具体机械结构的特性,可在不同机器人平台上迁移,为自主评估、修正自身尝试并持续改进提供预测基础。我们提出无需训练的动态引导动作修正(DGAC)方法:将每次失败视为查询,由学习到的模型生成并排序可能的纠正动作,使失败转化为下一轮策略更新的监督信号。在涵盖移动操作臂与固定机械臂的七个真实世界操作任务中,该方法使成功率从40%提升至81%,且适用于多种策略模型。结果表明,人类先验与机器人失败结合可实现可扩展的自主策略优化。

原文摘要 · Abstract (English)

A central question in robot learning is how to acquire skills from the kinds of data that humans learn from: passive observation, embodied practice, and the experience of failure. Human videos provide the first of these in abundance, and prior work has shown they can initialize useful policies. Far less clear is whether they can support the second and third: whether priors extracted from human videos can ground a robot's own attempts well enough to evaluate them, correct them, and improve from them. In this work, we show that human videos can be used to learn embodiment-agnostic action, dynamics, and value representations that transfer across robot embodiments, providing the predictive foundation required for robots to autonomously improve from their own rollouts and failures. We introduce Dynamics-Guided Action Correction (DGAC), a training-free approach that leverages these adapted models to repair failed states: each failure becomes a query for which the learned models propose and rank corrective actions, turning failures into supervision for the next policy update. Across seven real-world manipulation tasks spanning both a mobile manipulator and a static manipulator arm, our approach improves success rates from 40% to 81% across multiple policy backbones, demonstrating cross-embodiment robot self-improvement from human-video priors. These results show that human priors and robot failures can be combined to enable scalable autonomous policy improvement. Project page: https://ethz-mrl.github.io/robot-self-improvement-website/.

机器人学习自改进人类视频动作修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。