arXiv:2508.07642cs.AIcs.CL2025-08中稿 · ACL被引 4

将导航拆解为可解释的技能模块,提升复杂场景下的泛化能力。

Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents

  • 把导航分解为垂直移动、区域识别等原子技能,每项由专用代理处理。
  • 在GSA-R2R上实现最佳泛化性能,新指令和未见环境表现显著优于基线。
  • 无需人工标注,用合成数据+视觉语言模型路由,自动选择最优技能代理。

视觉-语言导航(VLN)要求智能体理解自然语言指令并在复杂3D环境中导航。尽管已有大规模预训练与数据增强推动进展,当前方法在面对需复杂空间和时间推理的新场景时仍难以泛化。本文提出SkillNav,一种基于技能的模块化框架,将导航任务分解为可解释的原子技能(如垂直移动、区域识别、停止暂停),每个技能由专门代理处理。为支持无标注的技能专项训练,构建了生成多样化、语言自然的技能特定指令-轨迹对的合成数据流水线。进一步引入无需训练的视觉-语言模型(VLM)路由机制,在每一步动态匹配子目标与视觉观察及历史动作,选择最合适的代理。SkillNav在主流基准上表现优异,并在包含新指令风格和未见环境的GSA-R2R上达到先进水平,显著提升泛化能力。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3D environments. While recent progress has been driven by large-scale pre-training and data augmentation, current methods still struggle to generalize to unseen scenarios, particularly when complex spatial and temporal reasoning is required. In this work, we propose SkillNav, a modular framework that introduces structured, skill-based reasoning into Transformer-based VLN agents. Our method decomposes navigation into a set of interpretable atomic skills (e.g., Vertical Movement, Area and Region Identification, Stop and Pause), each handled by a specialized agent. To support targeted skill training without manual data annotation, we construct a synthetic dataset pipeline that generates diverse, linguistically natural, skill-specific instruction-trajectory pairs. We then introduce a novel training-free Vision-Language Model (VLM)-based router, which dynamically selects the most suitable agent at each time step by aligning sub-goals with visual observations and historical actions. SkillNav obtains competitive results on commonly used benchmarks and establishes state-of-the-art generalization to the GSA-R2R, a benchmark with novel instruction styles and unseen environments.

视觉语言导航技能分解泛化能力模块化代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。