arXiv:2410.13002cs.ROcs.AI2024-10被引 3

用预训练视觉语言模型提取特征,实现小样本下跨场景的文本指令飞行导航。

Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features

  • 冻结VLM提取视觉语义特征,生成空间感知嵌入
  • 小样本模拟数据训练即可在真实场景中泛化
  • 适合需要快速适应新任务和指令的无人机应用

端到端学习直接将感官输入映射为动作,为复杂机器人任务构建高度集成高效的策略。然而,这类模型往往难以泛化到训练场景之外,限制了对新环境、新任务和新概念的适应能力。本文研究了在未见文本指令和视觉分布变化下,实现鲁棒闭环控制所需的最小数据量与架构调整。提出Flex(Fly lexically)框架,利用预训练视觉语言模型(VLMs)作为冻结的逐片特征提取器,生成融合语义与视觉信息的空间感知嵌入。在四旋翼飞向目标任务中,仅需少量模拟数据通过行为克隆训练的智能体,即可成功泛化至包含多种新目标和指令表述的真实场景。

原文摘要 · Abstract (English)

End-to-end learning directly maps sensory inputs to actions, creating highly integrated and efficient policies for complex robotics tasks. However, such models often struggle to generalize beyond their training scenarios, limiting adaptability to new environments, tasks, and concepts. In this work, we investigate the minimal data requirements and architectural adaptations necessary to achieve robust closed-loop performance with vision-based control policies under unseen text instructions and visual distribution shifts. Our findings are synthesized in Flex (Fly lexically), a framework that uses pre-trained Vision Language Models (VLMs) as frozen patch-wise feature extractors, generating spatially aware embeddings that integrate semantic and visual information. We demonstrate the effectiveness of this approach on a quadrotor fly-to-target task, where agents trained via behavior cloning on a small simulated dataset successfully generalize to real-world scenes with diverse novel goals and command formulations.

视觉导航多模态飞行控制小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。