提出SAME模型,让智能体用同一套系统完成多种语言导航任务。
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
- 用状态自适应专家混合机制,动态选择适合当前任务的决策模块。
- 在7种导航任务上表现媲美专用模型,实现通用性与性能兼顾。
- 适合研究多任务视觉导航、通用智能体的开发者参考。
语言引导的视觉导航研究通常分为高层类别搜索和低层语言导航两类,前者关注探索过程,后者聚焦于执行详细文本指令。尽管任务重点不同,但理解指令、感知环境与推理动作决策的核心需求一致。本文将多种导航任务统一到一个通用框架中,探究知识共享与任务特异性能力利用的挑战,提出一种状态自适应专家混合(State-Adaptive Mixture of Experts, SAME)模型,使智能体能根据不同粒度的语言指令与动态观测做出有效决策。基于SAME,我们构建了一个可同时应对七种导航任务的通用智能体,在性能上超越或媲美专用模型。
原文摘要 · Abstract (English)
The academic field of learning instruction-guided visual navigation can be generally categorized into high-level category-specific search and low-level language-guided navigation, depending on the granularity of language instruction, in which the former emphasizes the exploration process, while the latter concentrates on following detailed textual commands. Despite the differing focuses of these tasks, the underlying requirements of interpreting instructions, comprehending the surroundings, and inferring action decisions remain consistent. This paper consolidates diverse navigation tasks into a unified and generic framework -- we investigate the core difficulties of sharing general knowledge and exploiting task-specific capabilities in learning navigation and propose a novel State-Adaptive Mixture of Experts (SAME) model that effectively enables an agent to infer decisions based on different-granularity language and dynamic observations. Powered by SAME, we present a versatile agent capable of addressing seven navigation tasks simultaneously that outperforms or achieves highly comparable performance to task-specific agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。