用手机远程操控机器人,低成本收集高质量示范数据。
COBALT: Crowdsourcing Robot Learning via Cloud-Based Teleoperation with Smartphones

- 通过云端架构支持多用户并发操作,单块GPU可同时处理8人
- 用手机操作效果媲美专业设备,5天内跨9国收集超7500条数据
- 内置实时质量监控与训练课程,提升数据可用性,适合大规模数据采集
模仿学习在机器人操作中受限于高质量示范数据的稀缺。本文提出COBALT平台,通过向量化环境和负载均衡架构,实现单块GPU支持多达8名用户并发远程操控,端到端延迟低于100毫秒,帧率保持20赫兹。用户可使用手机、VR头显等常见设备接入。系统配备内存缓存与高效视频流,确保控制与渲染同步。实测支持256个模拟客户端跨8块GPU稳定运行。用户研究表明,手机操控表现不逊于专用硬件,且更便捷。平台自动记录实时指标筛选低质演示,并引入结构化培训课程提升数据质量。基于此,我们五日内以手机为工具,在九个国家收集了超过7500次示范(累计50+小时),并验证了该数据集可用于训练先进模仿学习模型。
原文摘要 · Abstract (English)
The scarcity of large-scale, high-quality demonstration data remains a bottleneck in scaling imitation learning for robotic manipulation. We present COBALT, a teleoperation platform designed to democratize robot learning at scale both in simulation and in the real world. By leveraging vectorized environments, our scalable, load-balanced infrastructure supports concurrent teleoperation by multiple users on a single GPU, yielding a significant reduction in teleoperation cost. Operators can connect from nearly anywhere on Earth using commonly available devices, including single or dual smartphones, VR headsets, 3D mice, and keyboards. An inmemory data cache and efficient video streaming keep control and rendering synchronous, sustaining dozens of concurrent users at 20 Hz with sub-100 ms end-to-end latency for up to 8 concurrent users per GPU. We also demonstrate stable operation supporting 256 simulated clients across 8 GPUs, underscoring the system's ability to scale across hardware and within individual servers. We perform a comprehensive user study showing that phone-based teleoperation performs comparably to or better than specialized hardware, enabling faster, more ergonomic data collection. To ensure data quality, COBALT logs a suite of real-time metrics to automatically filter suboptimal demonstrations. We further demonstrate that a structured user training curriculum significantly improves data collection quality. Guided by insights from our user study, we crowdsource the collection of a large-scale, high-quality pilot dataset with 7500+ demonstrations (50+ hours) collected with smartphones across nine countries over five days. We validate the dataset's quality by training state-of-the-art imitation learning algorithms. Please visit https://cobalt-teleop.github.io/ for more details.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。