SPOT让操作员在远程操控人形机器人时,能持续保持对环境的全面感知,解决长时间操作中的视觉盲区问题。
SPOT: Spatial Perception-Oriented Long-Horizon Humanoid Teleoperation

- 通过广角立体显示与视点解耦设计,实现稳定且可自由观察的机器人视角
- 在长时程任务中提升操作效率、准确率和恢复速度,显著减少失误
- 适合需要精细空间感知的人形机器人数据采集场景,如复杂交互与动态任务
高质量示范数据正成为训练通用人形机器人的关键瓶颈。尽管近期人形机器人遥操作系统在人体动作向机器人动作映射方面取得显著进展,但长时程的移动操纵仍需操作员维持任务相关的空间感知能力,例如物体位置、周围环境及机器人姿态。我们称这种感知范围为操作员的感知视野。然而现有方法常缩短这一视野:狭窄视角遗漏边缘事件,机器人搭载相机在运动中不稳定,且头部视角与机器人动作耦合导致环顾四周干扰机器人运动。我们提出SPOT,一种面向空间感知的虚拟现实遥操作系统,用于收集长时程人形机器人示范数据。SPOT结合机器人搭载双目鱼眼相机、宽视场立体显示器、视点解耦自由观察能力和视觉稳定技术,提供宽广、稳定且可主动检查的以机器人为中心的视角。不同于传统第一人称界面,SPOT将视觉探索与机器人执行解耦:第一人称立体观测被渲染在操作员周围的虚拟半球上,自然的头部转动仅改变观察方向而非控制机器人头部、相机或躯干。我们在涵盖跌倒恢复、边缘取物、大工作空间双臂操作、精细对齐和动态交互的感知关键任务中评估SPOT,结果表明其在效率、准确性和恢复速度上均有提升,证明了其在用户友好且可扩展的长时程人形机器人数据采集中的有效性。
原文摘要 · Abstract (English)
High-quality demonstration data is becoming a central bottleneck for training general-purpose humanoid robots. While recent humanoid teleoperation systems have made substantial progress in retargeting human motion to robot motion, long-horizon loco-manipulation requires another capability: operators must maintain task-relevant spatial awareness over time, e.g., object locations, surrounding environments, the robot's pose. We call the extent of this awareness the operator's perceptual horizon. However, existing methods often shorten this: narrow views miss peripheral events, robot-mounted cameras become unstable during locomotion, and coupled head-view control makes looking around interfere with robot motion. We present SPOT, a Spatial Perception-Oriented VR Teleoperation system for collecting long-horizon humanoid demonstration data by providing extended perceptual horizon. SPOT combines a robot-mounted binocular fisheye camera, a wide-field stereoscopic display, viewpoint-decoupled free-looking, and visual stabilization to provide a robot-centric view that is wide, stable, and actively inspectable. Unlike conventional egocentric interfaces, SPOT decouples visual exploration from robot actuation: the egocentric stereo observation is rendered on a virtual hemisphere around the operator, so natural head rotations change where the operator looks within the wide-field view rather than commanding the robot head, camera, or torso. We evaluate SPOT on perception-critical humanoid data-collection tasks spanning drop recovery, peripheral retrieval, large-workspace bimanual manipulation, fine alignment, and dynamic interaction. SPOT improves efficiency, accuracy, and recovery speed, demonstrating its effectiveness for user-friendly and scalable long-horizon humanoid data collection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。