arXiv:2502.13894cs.ROcs.CV2025-02ICRA被引 13

用视觉预测模型让机器人零样本导航,无需预先建图。

NavigateDiff: Visual Predictors are Zero-Shot Navigation Assistants

  • 用大模型+扩散网络构建视觉预测器,预判下一步看到的画面。
  • 引入时间历史信息使预测与真实场景对齐,提升导航准确性。
  • 适合需要快速适应新环境的家用机器人,尤其在仿真和真实场景均有效。

家庭机器人在陌生环境中导航面临巨大挑战,需识别并推理新装饰与布局。现有强化学习方法难以直接迁移至新环境,因依赖大量地图构建与探索,效率低下。为此,我们尝试将预训练基础模型的逻辑知识与泛化能力迁移至零样本导航任务。通过整合大型视觉语言模型与扩散网络,提出 mname ~ 方法,构建一个持续预测智能体下一步潜在观测的视觉预测器,从而辅助机器人生成稳健动作。为进一步适配导航的时间特性,引入时间历史信息,确保预测图像与导航场景一致。随后设计信息融合框架,将预测的未来帧作为引导嵌入目标达成策略中,解决下游图像导航任务。该方法显著提升了导航控制力与跨仿真与真实环境的泛化能力。大量实验验证了其鲁棒性与通用性,展现出在多样化场景中提升机器人导航效率与效果的巨大潜力。

原文摘要 · Abstract (English)

Navigating unfamiliar environments presents significant challenges for household robots, requiring the ability to recognize and reason about novel decoration and layout. Existing reinforcement learning methods cannot be directly transferred to new environments, as they typically rely on extensive mapping and exploration, leading to time-consuming and inefficient. To address these challenges, we try to transfer the logical knowledge and the generalization ability of pre-trained foundation models to zero-shot navigation. By integrating a large vision-language model with a diffusion network, our approach named \mname ~constructs a visual predictor that continuously predicts the agent's potential observations in the next step which can assist robots generate robust actions. Furthermore, to adapt the temporal property of navigation, we introduce temporal historical information to ensure that the predicted image is aligned with the navigation scene. We then carefully designed an information fusion framework that embeds the predicted future frames as guidance into goal-reaching policy to solve downstream image navigation tasks. This approach enhances navigation control and generalization across both simulated and real-world environments. Through extensive experimentation, we demonstrate the robustness and versatility of our method, showcasing its potential to improve the efficiency and effectiveness of robotic navigation in diverse settings.

零样本导航视觉预测扩散模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。