arXiv:2604.00416cs.ROcs.AI2026-04

仅用5小时人类行走数据,让机器人学会在新环境中自主导航。

Learning Humanoid Navigation from Human Data

  • 用扩散模型预测未来路径分布,结合360°视觉记忆和深层特征。
  • 10步去噪实现实时推理,在多种场景下零样本部署成功。
  • 无需训练即可自然生成避障、等门、绕人群等智能行为。

我们提出EgoNav系统,通过仅5小时人类行走数据,使拟人机器人在未见过的室内外环境中自主导航,无需任何机器人数据或微调。系统采用扩散模型根据历史轨迹预测未来路径分布,融合360°彩色、深度与语义信息的视觉记忆,并利用冻结的DINOv3骨干网络提取深度传感器无法感知的外观特征。通过混合采样策略,实现在10步去噪中完成实时推理,再由滚动时域控制器从预测分布中选择路径。离线评估显示其在避障和多模态覆盖方面优于基线方法;零样本部署于Unitree G1拟人机器人,成功穿越多种未知环境。学习到的先验自然催生了等待开门、绕行人群、避开玻璃墙等行为。数据集与训练模型将公开。官网:https://egonav.weizhuowang.com

原文摘要 · Abstract (English)

We present EgoNav, a system that enables a humanoid robot to traverse diverse, unseen environments by learning entirely from 5 hours of human walking data, with no robot data or finetuning. A diffusion model predicts distributions of plausible future trajectories conditioned on past trajectory, a 360 deg visual memory fusing color, depth, and semantics, and video features from a frozen DINOv3 backbone that capture appearance cues invisible to depth sensors. A hybrid sampling scheme achieves real-time inference in 10 denoising steps, and a receding-horizon controller selects paths from the predicted distribution. We validate EgoNav through offline evaluations, where it outperforms baselines in collision avoidance and multi-modal coverage, and through zero-shot deployment on a Unitree G1 humanoid across unseen indoor and outdoor environments. Behaviors such as waiting for doors to open, navigating around crowds, and avoiding glass walls emerge naturally from the learned prior. We will release the dataset and trained models. Our website: https://egonav.weizhuowang.com

机器人导航扩散模型人类数据零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。