用人类失败恢复动作训练机器人,效率比机器人操作高10倍以上。
EgoRecovery: Acquiring Failure Recovery Ability Through Human Recovery Demonstration

- 用人类视角视频记录恢复片段,大幅提升数据采集效率。
- 仅需少量机器人恢复数据即可实现从人类意图到机器人动作的映射。
- 部署时自动判断何时需纠正,显著提升真实场景任务成功率。
可靠的具身机器人需具备从失败中恢复并重试任务的能力,以在非结构化、嘈杂的真实环境中稳定运行。这需要基于包含恢复行为的数据训练策略。然而,通过机器人遥操作收集此类数据难以扩展,因诱发多样失败状态、执行修正动作及重置环境耗时较长。失败模式多样性进一步加剧了对恢复数据的需求,所需量远超成功示范。本文提出,利用捕捉失败恢复过程的共视人类数据可作为可扩展替代方案。通过高效配置任务级失败情境并录制短时恢复片段,人类操作员每小时生成的有效恢复数据量超过机器人遥操作的10倍。为解决人机间表征差距,提出EgoRecovery联合训练框架,将人类恢复示范对齐至与机器人数据共享的紧凑校正意图空间,该空间捕捉修正的时间与幅度特征。仅需少量机器人恢复示范,即可将此意图映射为可执行的机器人动作。部署时,学习的恢复门控模块根据机器人观测判断是否需要纠正,并仅在恢复状态下激活校正意图。真实世界恢复任务实验表明,EgoRecovery在失败起点上的成功率优于纯机器人恢复、直接联合训练人类数据及直接意图迁移基线。
原文摘要 · Abstract (English)
Robust embodied robots should be able to recover from failures and retry tasks in order to operate reliably in unstructured and noisy real-world environments. Achieving this capability requires training policies on data that captures recovery behaviors. However, collecting such data through robot teleoperation is difficult to scale, as it is time-consuming to induce diverse failure states, perform corrective actions, and reset the environment. This challenge is further exacerbated by the high diversity of failure modes, which demands substantially more recovery data than success demonstrations. In this work, we show that egocentric human data capturing failure recovery processes provides a scalable alternative. By efficiently arranging task-level failure configurations and recording short recovery segments, human operators can generate more than 10x as much valid recovery data per hour compared to robot teleoperation under our protocol. To address the embodiment gap between human and robot, we propose EgoRecovery, a co-training framework for learning recovery behavior, where human recovery demonstrations are aligned to a compact corrective-intent space shared with robot data, which captures the timing and magnitude of correction. Only a small number of robot recovery demonstrations are required to connect this intent to executable robot actions. At deployment, a learned recovery gate predicts when correction is needed from robot observations and activates the corrective intent only in recovery states. Experiments on real-world recovery tasks show that EgoRecovery improves success from failure starts over robot-only recovery, direct co-training with human recovery data, and direct intent-transfer baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。