用视觉边界做导航锚点,让机器人零样本适配复杂环境。
OpenFrontier: General Navigation with Visual-Language Grounded Frontiers
- 以视觉边界作为语义锚点,实现稀疏子目标定位与导航
- 无需任务微调,在多个基准上实现强零样本表现
- 轻量级设计,适合真实机器人部署,兼容多种视觉语言模型
开放世界导航要求机器人在复杂的日常环境中做出决策,并适应灵活的任务需求。传统导航方法通常依赖密集的3D重建和手工设计的目标度量,限制了其在不同任务和环境中的泛化能力。近年来,视觉语言导航(VLN)和视觉语言动作(VLA)模型实现了基于自然语言的端到端策略,但通常需要交互式训练、大规模数据收集或移动代理的任务特定微调。本文将导航建模为稀疏子目标识别与到达问题,发现为高层语义先验提供视觉锚点可实现高效的目标条件导航。基于此洞察,我们选择视觉前沿作为语义锚点,提出OpenFrontier框架:该框架无需任务特定训练或微调,可无缝集成多种视觉语言先验模型。OpenFrontier采用轻量级系统设计,无需密集3D语义映射、任务特定策略训练或模型微调,可在多个导航基准上实现强零样本性能,并成功部署于真实移动机器人。
原文摘要 · Abstract (English)
Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Conventional navigation approaches often rely on dense 3D reconstruction and hand-crafted goal metrics, which limits their generalization across tasks and environments. Recent advances in vision-language navigation (VLN) and vision-language-action (VLA) models enable end-to-end policies conditioned on natural language, but typically require interactive training, large-scale data collection, or task-specific fine-tuning with a mobile agent. We formulate navigation as a sparse subgoal identification and reaching problem and observe that providing visual anchoring targets for high-level semantic priors enables highly efficient goal-conditioned navigation. Based on this insight, we select visual frontiers as semantic anchors and propose OpenFrontier, a navigation framework that requires no task-specific training or fine-tuning and seamlessly integrates diverse vision-language prior models. OpenFrontier enables efficient navigation with a lightweight system design, without dense 3D semantic mapping, task-specific policy training, or model fine-tuning. We evaluate OpenFrontier across multiple navigation benchmarks and demonstrate strong zero-shot performance, as well as effective real-world deployment on a mobile robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。