arXiv:2602.01930cs.ROcs.CV2026-02被引 2

用视觉语言模型让机器人在未知环境里按语义指令自主探索

LIEREx: Language-Image Embeddings for Robotic Exploration

  • 结合视觉语言大模型与3D语义场景图,实现开放集语义建图
  • 支持按自然语言指令在部分未知环境中定位目标物体
  • 适合需要灵活理解环境的自主机器人应用

语义地图使机器人能够推理周围环境,完成导航、找物和探索未知区域等任务。传统方法虽能提供精确的几何表示,但受限于预设的符号词汇表,难以处理设计时未定义的新知识。近年来,如CLIP等视觉语言基础模型(VLFMs)实现了开放集映射,将物体编码为高维嵌入而非固定标签。本文提出的LIEREx将这些VLFMs与成熟的3D语义场景图结合,使自主代理能在部分未知环境中根据目标指令进行探索。

原文摘要 · Abstract (English)

Semantic maps allow a robot to reason about its surroundings to fulfill tasks such as navigating known environments, finding specific objects, and exploring unmapped areas. Traditional mapping approaches provide accurate geometric representations but are often constrained by pre-designed symbolic vocabularies. The reliance on fixed object classes makes it impractical to handle out-of-distribution knowledge not defined at design time. Recent advances in Vision-Language Foundation Models, such as CLIP, enable open-set mapping, where objects are encoded as high-dimensional embeddings rather than fixed labels. In LIEREx, we integrate these VLFMs with established 3D Semantic Scene Graphs to enable target-directed exploration by an autonomous agent in partially unknown environments.

机器人探索视觉语言模型语义地图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。