arXiv:2502.14195cs.CVcs.RO2025-02被引 7

用文字描述匹配图像,实现基于文本的精准定位。

Bridging Text and Vision: A Multi-View Text-Vision Registration Approach for Cross-Modal Place Recognition

  • 通过360度视觉与T5模型结合,将文本描述映射到图像特征。
  • 在街景数据集上达到57%的顶1准确率,5米内92%顶10准确率。
  • 适合需要语言理解的机器人导航、智能配送等场景。

移动机器人需具备先进的自然语言理解能力,以准确识别位置并完成如包裹递送等任务。然而,传统视觉定位(VPR)方法仅依赖单视角视觉信息,无法解析人类语言描述。为解决此问题,本文提出一种多视角(360°全景)文本-视觉配准方法Text4VPR,首次实现仅使用文本描述匹配图像数据库。Text4VPR采用冻结的T5语言模型提取全局文本嵌入,并利用带温度系数的Sinkhorn算法将局部词元分配至对应聚类,聚合图像视觉描述符。训练阶段强调文本-图像对间的精准对齐;推理阶段采用级联交叉注意力余弦对齐(CCCA)处理文本与图像组之间的内部错位,进而基于文本-图像组描述进行精确位置匹配。在自建的首个文本到图像的VPR数据集Street360Loc上,Text4VPR建立了稳健基线,测试集上5米半径内取得57%的顶1准确率和92%的顶10准确率,表明从文本描述到图像的定位不仅可行,且具有显著提升潜力。

原文摘要 · Abstract (English)

Mobile robots necessitate advanced natural language understanding capabilities to accurately identify locations and perform tasks such as package delivery. However, traditional visual place recognition (VPR) methods rely solely on single-view visual information and cannot interpret human language descriptions. To overcome this challenge, we bridge text and vision by proposing a multiview (360° views of the surroundings) text-vision registration approach called Text4VPR for place recognition task, which is the first method that exclusively utilizes textual descriptions to match a database of images. Text4VPR employs the frozen T5 language model to extract global textual embeddings. Additionally, it utilizes the Sinkhorn algorithm with temperature coefficient to assign local tokens to their respective clusters, thereby aggregating visual descriptors from images. During the training stage, Text4VPR emphasizes the alignment between individual text-image pairs for precise textual description. In the inference stage, Text4VPR uses the Cascaded Cross-Attention Cosine Alignment (CCCA) to address the internal mismatch between text and image groups. Subsequently, Text4VPR performs precisely place match based on the descriptions of text-image groups. On Street360Loc, the first text to image VPR dataset we created, Text4VPR builds a robust baseline, achieving a leading top-1 accuracy of 57% and a leading top-10 accuracy of 92% within a 5-meter radius on the test set, which indicates that localization from textual descriptions to images is not only feasible but also holds significant potential for further advancement, as shown in Figure 1.

跨模态定位文本视觉融合机器人导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。