arXiv:2510.19655cs.RO2025-10被引 6

LaViRA让机器人零样本导航,靠分层语言视觉动作解耦提升泛化能力。

LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments

  • 分三层:语言规划、视觉定位、机器人控制,各用不同大模型优势
  • 在未见环境中导航成功率显著高于现有方法,突破零样本限制
  • 适合需要强泛化和可解释性的现实机器人导航场景

LaViRA 是一种用于连续环境中的零样本视觉语言导航(VLN-CE)的框架,要求智能体在未见过的环境中仅凭自然语言指令进行导航而无需预先训练。现有方法面临关键权衡:要么依赖环境特定的路径点预测器,导致场景泛化能力差;要么未能充分利用大模型的推理能力。本文提出 LaViRA,通过将动作分解为粗到细的层次结构——语言动作(高阶规划)、视觉动作(中阶感知对齐)、机器人动作(低阶控制),使多模态大语言模型(MLLM)在各阶段发挥其独特优势。该模块化设计实现了强大推理、精准感知与实际控制的统一,在 VLN-CE 基准上显著超越现有最先进方法,展现出卓越的未见环境泛化能力,同时保持部署透明性与高效性。

原文摘要 · Abstract (English)

LaViRA: Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to navigate unseen environments based on natural language instructions without any prior training. Current methods face a critical trade-off: either rely on environment-specific waypoint predictors that limit scene generalization, or underutilize the reasoning capabilities of large models during navigation. We introduce LaViRA, a simple yet effective zero-shot framework that addresses this dilemma by decomposing action into a coarse-to-fine hierarchy: Language Action for high-level planning, Vision Action for middle-level perceptual grounding, and Robot Action for low-level control. This modular decomposition allows us to leverage the distinct strengths of different scales of Multimodal Large Language Models (MLLMs) at each stage, creating a system that is powerful in its reasoning, grounding and practical control. LaViRA significantly outperforms existing state-of-the-art methods on the VLN-CE benchmark, demonstrating superior generalization capabilities in unseen environments, while maintaining transparency and efficiency for real-world deployment. Project page: https://robo-lavira.github.io/lavira-zs-vln/

视觉导航零样本多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。