arXiv:2501.03304cs.ROcs.LG2025-01被引 1

让机器人地图会说话:用语言描述环境,支持动态更新与多视角一致

LiLMaps: Learnable Implicit Language Maps

  • 用视觉-语言特征融合优化隐式地图的解码器,支持新物体自动加入
  • 解决不同角度观察时语言描述不一致的问题,提升地图语义一致性
  • 适合做智能机器人场景理解与自然人机交互的研究者参考

当前机器人领域的一个趋势是利用大语言模型(LLMs)实现非预定义指令执行和自然的人机交互。拥有带语言表示的环境地图,有助于LLMs进一步理解与操作场景。这种综合场景表征可为自主机器人提供多种交互方式。本文提出LiLMaps,通过融合视觉-语言特征增强增量隐式地图构建。具体包括:(i) 提出一种隐式语言地图的解码器优化技术,可在新物体出现时有效更新;(ii) 解决不同视角下视觉-语言预测不一致的问题。实验表明,LiLMaps显著提升了性能。

原文摘要 · Abstract (English)

One of the current trends in robotics is to employ large language models (LLMs) to provide non-predefined command execution and natural human-robot interaction. It is useful to have an environment map together with its language representation, which can be further utilized by LLMs. Such a comprehensive scene representation enables numerous ways of interaction with the map for autonomously operating robots. In this work, we present an approach that enhances incremental implicit mapping through the integration of vision-language features. Specifically, we (i) propose a decoder optimization technique for implicit language maps which can be used when new objects appear on the scene, and (ii) address the problem of inconsistent vision-language predictions between different viewing positions. Our experiments demonstrate the effectiveness of LiLMaps and solid improvements in performance.

机器人语言地图视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。