arXiv:2603.07624cs.RO2026-03

用视觉模型的3D几何先验,让纯摄像头机器人在仿真中学会稳定行走。

GeoLoco: Leveraging 3D Geometric Priors from Visual Foundation Model for Robust RGB-Only Humanoid Locomotion

  • 利用冻结的视觉基础模型提取单目图像的3D隐空间表示。
  • 在仿真中训练后,直接部署到真实机器人上实现零样本迁移。
  • 通过双头辅助学习防止纹理过拟合,提升地形适应能力。

当前感知式人形机器人行走主要依赖主动深度传感器,但这种以深度为中心的方法忽略了视觉世界的丰富语义与密集外观线索,使底层控制与高层推理脱节。虽然单目RGB提供普遍且信息密集的替代方案,但直接从原始2D像素进行端到端强化学习存在极端样本效率低下和严重的仿真到现实崩溃问题,根源在于几何尺度信息的丢失。为打破僵局,我们提出GeoLoco,一种完全基于RGB的行走框架,将单目图像视为高维3D潜在表示,借助冻结的、具备尺度感知能力的视觉基础模型(VFM)的强大几何先验。不同于简单的特征拼接,我们设计了本体感觉查询的多头交叉注意力机制,动态关注随机器人实时步态阶段变化的任务关键拓扑特征。关键的是,为防止策略对表面纹理过拟合,引入双头辅助学习方案,显式正则化使高维潜在空间严格对齐物理地形几何,确保鲁棒的零样本仿真到现实迁移。仅在仿真中训练的GeoLoco成功实现对Unitree G1人形机器人的稳健零样本迁移,并能有效应对复杂地形。

原文摘要 · Abstract (English)

The prevailing paradigm of perceptive humanoid locomotion relies heavily on active depth sensors. However, this depth-centric approach fundamentally discards the rich semantic and dense appearance cues of the visual world, severing low-level control from the high-level reasoning essential for general embodied intelligence. While monocular RGB offers a ubiquitous, information-dense alternative, end-to-end reinforcement learning from raw 2D pixels suffers from extreme sample inefficiency and catastrophic sim-to-real collapse due to the inherent loss of geometric scale. To break this deadlock, we propose GeoLoco, a purely RGB-driven locomotion framework that conceptualizes monocular images as high-dimensional 3D latent representations by harnessing the powerful geometric priors of a frozen, scale-aware Visual Foundation Model (VFM). Rather than naive feature concatenation, we design a proprioceptive-query multi-head cross-attention mechanism that dynamically attends to task-critical topological features conditioned on the robot's real-time gait phase. Crucially, to prevent the policy from overfitting to superficial textures, we introduce a dual-head auxiliary learning scheme. This explicit regularization forces the high-dimensional latent space to strictly align with the physical terrain geometry, ensuring robust zero-shot sim-to-real transfer. Trained exclusively in simulation, GeoLoco achieves robust zero-shot transfer to the Unitree G1 humanoid and successfully negotiates challenging terrains.

人形机器人纯视觉几何先验零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。