用手机看机器人预测轨迹,实时优化策略,无需真机器人。
RoboPocket: Improve Robot Policies Instantly with Your Phone
- 手机通过AR预览机器人动作预测,指导数据收集
- 相比离线方法,数据效率提升一倍,分钟级完成迭代
- 适合希望快速优化机器人策略的研究者和开发者
模仿学习的扩展受限于数据采集效率。手持设备虽可实现大规模野外数据采集,但多为开环操作:操作者无法感知策略弱点,导致关键状态分布覆盖不足。而交互式方法如DAgger虽能缓解协变量偏移,却依赖真实机器人执行,成本高且难扩展。为此,我们提出RoboPocket,一种基于单部消费级智能手机的便携式系统,实现无需机器人的即时策略迭代。其核心是远程推理框架,通过增强现实(AR)视觉前瞻可视化策略预测轨迹,使采集者可主动识别潜在失败并聚焦于策略薄弱区域。此外,我们设计了异步在线微调流水线,持续用新数据更新策略,实现分钟级闭环学习。大量实验表明,RoboPocket符合数据缩放定律,相比离线策略数据效率提升2倍,突破长期效率瓶颈;在分布式环境下,每人少量交互修正即可提升样本效率达2倍。项目主页与视频:https://robo-pocket.github.io。
原文摘要 · Abstract (English)
Scaling imitation learning is fundamentally constrained by the efficiency of data collection. While handheld interfaces have emerged as a scalable solution for in-the-wild data acquisition, they predominantly operate in an open-loop manner: operators blindly collect demonstrations without knowing the underlying policy's weaknesses, leading to inefficient coverage of critical state distributions. Conversely, interactive methods like DAgger effectively address covariate shift but rely on physical robot execution, which is costly and difficult to scale. To reconcile this trade-off, we introduce RoboPocket, a portable system that enables Robot-Free Instant Policy Iteration using single consumer smartphones. Its core innovation is a Remote Inference framework that visualizes the policy's predicted trajectory via Augmented Reality (AR) Visual Foresight. This immersive feedback allows collectors to proactively identify potential failures and focus data collection on the policy's weak regions without requiring a physical robot. Furthermore, we implement an asynchronous Online Finetuning pipeline that continuously updates the policy with incoming data, effectively closing the learning loop in minutes. Extensive experiments demonstrate that RoboPocket adheres to data scaling laws and doubles the data efficiency compared to offline scaling strategies, overcoming their long-standing efficiency bottleneck. Moreover, our instant iteration loop also boosts sample efficiency by up to 2$\times$ in distributed environments a small number of interactive corrections per person. Project page and videos: https://robo-pocket.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。