arXiv:2410.00388cs.RO2024-10被引 9

用视觉语言模型实现多目标搜索,导航更准更快

Find Everything: A General Vision Language Model Approach to Multi-Object Search

  • 通过多通道得分图同时追踪多个目标位置
  • 在模拟和真实环境中均优于现有方法
  • 适合需要高效寻物的机器人应用

多目标搜索(MOS)问题要求在不同环境中导航至一系列位置,以最大化找到目标物体的概率,同时最小化移动成本。本文提出一种新方法Finder,利用视觉语言模型(VLMs)在多样环境中定位多个物体。该方法引入多通道得分图,在导航过程中同步跟踪和推理多个目标,并采用结合场景级与物体级语义关联的得分图技术。在模拟和真实环境中的实验表明,Finder在性能上优于基于深度强化学习和VLMs的现有方法。消融研究验证了设计选择的有效性,可扩展性研究则证明其在目标数量增加时仍保持鲁棒性。

原文摘要 · Abstract (English)

The Multi-Object Search (MOS) problem involves navigating to a sequence of locations to maximize the likelihood of finding target objects while minimizing travel costs. In this paper, we introduce a novel approach to the MOS problem, called Finder, which leverages vision language models (VLMs) to locate multiple objects across diverse environments. Specifically, our approach introduces multi-channel score maps to track and reason about multiple objects simultaneously during navigation, along with a score map technique that combines scene-level and object-level semantic correlations. Experiments in both simulated and real-world settings showed that Finder outperforms existing methods using deep reinforcement learning and VLMs. Ablation and scalability studies further validated our design choices and robustness with increasing numbers of target objects, respectively. Website: https://find-all-my-things.github.io/

多目标搜索视觉语言模型机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。