不微调VLM也能提升视觉语言导航的物体识别能力
Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
- 用弱监督部分对比学习融合预训练VLM知识
- 在多个基准上超越基线,提升导航准确率
- 无需微调,兼顾效果与计算效率,适合资源受限场景
视觉语言导航(VLN)是具身智能中的基础任务,要求智能体根据自然语言指令在复杂环境中导航。现有方法存在三大挑战:依赖预训练视觉主干模型,在动态视角下表现不佳;使用未微调的大语言模型或视觉语言模型时,因缺乏领域知识而性能受限;微调虽能提升效果,但计算开销较大。为此,我们提出弱监督部分对比学习(WPCL),通过在感知过程中有效融入预训练视觉语言模型知识,提升智能体在动态视角下的物体识别能力,且无需对VLM进行微调。该方法增强了智能体对环境线索的理解与响应能力,同时保持高计算效率。实验表明,该方法在多个基准上均优于基线,验证了其有效性、鲁棒性和泛化能力。
原文摘要 · Abstract (English)
Visual Language Navigation (VLN) is a fundamental task within the field of Embodied AI, focusing on the ability of agents to navigate complex environments based on natural language instructions. Despite the progress made by existing methods, these methods often present some common challenges. First, they rely on pre-trained backbone models for visual perception, which struggle with the dynamic viewpoints in VLN scenarios. Second, the performance is limited when using pre-trained LLMs or VLMs without fine-tuning, due to the absence of VLN domain knowledge. Third, while fine-tuning LLMs and VLMs can improve results, their computational costs are higher than those without fine-tuning. To address these limitations, we propose Weakly-supervised Partial Contrastive Learning (WPCL), a method that enhances an agent's ability to identify objects from dynamic viewpoints in VLN scenarios by effectively integrating pre-trained VLM knowledge into the perception process, without requiring VLM fine-tuning. Our method enhances the agent's ability to interpret and respond to environmental cues while ensuring computational efficiency. Experimental results have shown that our method outperforms the baseline methods on multiple benchmarks, which validate the effectiveness, robustness and generalizability of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。