用人类回溯修正指导,让机器人低成本高效学技能
Genie Centurion: Accelerating Scalable Real-World Robot Training with Human Rewind-and-Refine Guidance
- 机器人出错时可回退到之前状态,由人提供纠正示范
- 任务成功率比现有方法高40%,用一半数据达到相近效果
- 适合多机器人协同训练,降低人力成本
尽管视觉-语言-动作(VLA)模型在多种任务中表现出强泛化能力,但实际部署机器人策略仍需大规模高质量的人类专家示范。然而,通过人类远程操控收集数据需持续专注,成本高且难扩展。为此,我们提出Genie Centurion(GCENT),一种基于人类回溯与修正指导的可扩展通用数据收集范式,支持机器人在部署中的交互式学习。GCENT从一个不完善的策略开始,随时间迭代优化。当机器人执行失败时,系统通过回退机制恢复到先前状态,由远程操作员提供纠正示范以改进策略。该框架结合任务哨兵模块,实现一人对多机器人的监督模式,能自主预测任务成功概率并在必要时请求人工干预。实证结果表明,GCENT在长时序、高精度任务中,任务成功率最高提升40%,且使用数据量不足现有方法的一半即可达到相当性能。我们还量化了多机器人场景下的数据产出效率,验证了其在真实环境中的可扩展性与成本效益。
原文摘要 · Abstract (English)
While Vision-Language-Action (VLA) models show strong generalizability in various tasks, real-world deployment of robotic policy still requires large-scale, high-quality human expert demonstrations. However, data collection via human teleoperation requires continuous operator attention, which is costly, hard to scale. To address this, we propose Genie Centurion (GCENT), a scalable and general data collection paradigm based on human rewind-and-refine guidance, enabling robots' interactive learning in deployment. GCENT starts at an imperfect policy and improves over time. When the robot execution failures occur, GCENT allows robots to revert to a previous state with a rewind mechanism, after which a teleoperator provides corrective demonstrations to refine the policy. This framework supports a one-human-to-many-robots supervision scheme with a Task Sentinel module, which autonomously predicts task success and solicits human intervention when necessary. Empirical results show that GCENT achieves up to 40% higher task success rates than state-of-the-art data collection methods, and reaches comparable performance using less than half the data in long-horizon and precise tasks. We also quantify the data yield-to-effort ratio under multi-robot scenarios, demonstrating GCENT's potential for scalable and cost-efficient robot policy training in real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。