用自车视角视觉实现四足机器人自主抓取,无需外部设备
SigLoMa: Learning Open-World Quadrupedal Loco-Manipulation from Ego-Centric Vision
- 用轻量几何表示法实现高可扩展的感知
- 5Hz检测下达成接近人类操作的动态抓取性能
- 适合无外部定位系统的实时机器人开发
设计开放世界的四足移动-操作系统极具挑战。传统基于外视觉的强化学习常面临样本效率极低和仿真到现实的巨大差距,且视觉跟踪固有的延迟与高频率浮动基控制需求根本冲突。因此现有系统严重依赖昂贵的外部运动捕捉和离线计算。为消除这些依赖,我们提出SigLoMa,一种完全在机、基于自车视角视觉的抓放框架。核心是引入Sigma Points,一种轻量级几何表征,保证高可扩展性与天然的仿真-现实对齐。为弥合感知慢与控制快的频率差异,设计了自车视角卡尔曼滤波器,实现鲁棒的高频状态估计。在学习方面,通过由提示位姿引导的主动采样课程缓解样本效率问题,并结合时间编码与模拟随机游走漂移解决机器人结构导致的视觉盲区。真实世界实验表明,仅使用5Hz(200ms延迟)的开放词汇检测器,SigLoMa成功完成多项动态移动-操作任务,性能接近专家级人类遥操作。
原文摘要 · Abstract (English)
Designing an open-world quadrupedal loco-manipulation system is highly challenging. Traditional reinforcement learning frameworks utilizing exteroception often suffer from extreme sample inefficiency and massive sim-to-real gaps. Furthermore, the inherent latency of visual tracking fundamentally conflicts with the high-frequency demands of precise floating-base control. Consequently, existing systems lean heavily on expensive external motion capture and off-board computation. To eliminate these dependencies, we present SigLoMa, a fully onboard, ego-centric vision-based pick-and-place framework. At the core of SigLoMa is the introduction of Sigma Points, a lightweight geometric representation for exteroception that guarantees high scalability and native sim-to-real alignment. To bridge the frequency divide between slow perception and fast control, we design an ego-centric Kalman Filter to provide robust, high-rate state estimation. On the learning front, we alleviate sample inefficiency via an Active Sampling Curriculum guided by Hint Poses, and tackle the robot's structural visual blind spots using temporal encoding coupled with simulated random-walk drift. Real-world experiments validate that, relying solely on a 5Hz (200 ms latency) open-vocabulary detector, SigLoMa successfully executes dynamic loco-manipulation across multiple tasks, achieving performance comparable to expert human teleoperation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。