arXiv:2509.16445cs.RO2025-09被引 11

直接微调视觉语言模型,实现高效通用的语义导航。

FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning

  • 用模拟环境数据微调VLM,直接根据视觉轨迹和目标选择探索方向。
  • 在HM3D ObjectNav上刷新SPL与成功率,对未见物体类别也表现优异。
  • 适合需要强泛化能力的机器人导航研究者参考。

让机器人助手在复杂环境中通过自由语言描述定位物体,是实际部署的关键能力。尽管基础模型(尤其是视觉语言模型,VLM)具备强大的语义理解能力,但如何有效将其大规模网络知识应用于具身决策仍是一大挑战。本文提出FiLM-Nav(用于导航的微调语言模型),直接将预训练的VLM作为导航策略。与仅使用零样本推理或用于地图标注的方法不同,FiLM-Nav基于原始视觉轨迹历史和导航目标,学习选择下一个最佳探索前沿。通过在包含ObjectNav、OVON、ImageNav及辅助空间推理任务的多样化模拟具身数据上进行微调,使VLM的预训练表征与目标驱动导航的动态和视觉模式紧密结合。该方法在开放词汇导航中达到新的SPL与成功率达到新高,在挑战性的HM3D-OVON基准上创下新的SPL记录,展现出对未见物体类别的强泛化能力。结果验证了在多样模拟具身数据上直接微调VLM,是实现高效且可泛化的语义导航的有效路径。

原文摘要 · Abstract (English)

Enabling robotic assistants to navigate complex environments and locate objects described in free-form language is a critical capability for real-world deployment. While foundation models, particularly Vision-Language Models (VLMs), offer powerful semantic understanding, effectively adapting their web-scale knowledge for embodied decision-making remains a key challenge. We present FiLM-Nav (Fine-tuned Language Model for Navigation), an approach that directly fine-tunes pre-trained VLM as the navigation policy. In contrast to methods that use foundation models primarily in a zero-shot manner or for map annotation, FiLM-Nav learns to select the next best exploration frontier by conditioning directly on raw visual trajectory history and the navigation goal. Leveraging targeted simulated embodied experience allows the VLM to ground its powerful pre-trained representations in the specific dynamics and visual patterns relevant to goal-driven navigation. Critically, fine-tuning on a diverse data mixture combining ObjectNav, OVON, ImageNav, and an auxiliary spatial reasoning task proves essential for achieving robustness and broad generalization. FiLM-Nav sets a new state-of-the-art in both SPL and success rate on HM3D ObjectNav among open-vocabulary methods, and sets a state-of-the-art SPL on the challenging HM3D-OVON benchmark, demonstrating strong generalization to unseen object categories. Our work validates that directly fine-tuning VLMs on diverse simulated embodied data is a highly effective pathway towards generalizable and efficient semantic navigation capabilities.

机器人导航视觉语言模型具身智能语义导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。