arXiv:2507.07299cs.ROcs.CV2025-07被引 1

提出多层特征图,提升零样本语义导航中的语言理解能力

MLFM: Multi-Layered Feature Maps for Richer Language Understanding in Zero-Shot Semantic Navigation

  • 构建可查询的多层语义地图,融合视觉语言特征
  • 在自然语言目标下实现更精准的属性与空间关系推理
  • 适合研究零样本语义导航与语言-视觉对齐的学者

近期大型视觉-语言模型推动了基于语言的语义导航进展,但缺乏清晰、以语言为核心的评估框架来检验智能体如何将指令中的词汇与环境对应。为此,我们提出LangNav——一个开放词汇、多物体导航数据集,包含自然语言目标描述(如‘去桌子上的红色短蜡烛’)及细粒度语言标注(如属性:颜色=红色,大小=短;关系:支撑=在……上)。这些标注支持系统性评估语言理解能力。为在此设置下评估,我们将多物体导航任务扩展为语言引导的多物体导航(LaMoN),要求智能体按语言指令依次找到多个目标。同时,我们提出多层特征图(MLFM)方法,从预训练视觉-语言特征构建可查询的多层语义地图,并证明其在推理目标描述中细粒度属性与空间关系上的有效性。在LangNav上的实验表明,MLFM优于当前最先进的零样本映射导航基线。

原文摘要 · Abstract (English)

Recent progress in large vision-language models has driven improvements in language-based semantic navigation, where an embodied agent must reach a target object described in natural language. Yet we still lack a clear, language-focused evaluation framework to test how well agents ground the words in their instructions. We address this gap by proposing LangNav, an open-vocabulary multi-object navigation dataset with natural language goal descriptions (e.g. 'go to the red short candle on the table') and corresponding fine-grained linguistic annotations (e.g., attributes: color=red, size=short; relations: support=on). These labels enable systematic evaluation of language understanding. To evaluate on this setting, we extend multi-object navigation task setting to Language-guided Multi-Object Navigation (LaMoN), where the agent must find a sequence of goals specified using language. Furthermore, we propose Multi-Layered Feature Map (MLFM), a novel method that builds a queryable, multi-layered semantic map from pretrained vision-language features and proves effective for reasoning over fine-grained attributes and spatial relations in goal descriptions. Experiments on LangNav show that MLFM outperforms state-of-the-art zero-shot mapping-based navigation baselines.

语义导航视觉语言多层特征零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。