arXiv:2604.18271cs.RO2026-04

让机器人用轻量图结构高效存取环境记忆,响应更快更自然。

EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents

论文配图:EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
图 1 · 摘自论文原文
  • 用轻量图+检索融合构建环境记忆,兼顾细节与高层语义。
  • 在NaVQA上推理和查询速度领先,准确率接近顶尖水平。
  • 可在真实机器人本地运行,适合需快速响应的交互场景。

随着智能体在机器人领域的应用发展,具备高效构建与检索记忆能力的机器人需求日益增长。在复杂环境中运行的机器人需建立记忆结构,以通过当前操作情境的表征实现有用的人机交互。人类与机器人互动时,期望其能提供关于位置、事件或物体的信息,这就要求机器人能在类人推理时间内给出精确回答,以显得响应迅速。我们提出EmbodiedLGR-Agent,一种由视觉语言模型(VLM)驱动的机器人智能体架构,可构建密集且高效的环境表示。该架构通过参数高效的VLM,将物体及其位置的低层信息存储在语义图中,同时利用传统检索增强架构保留观察场景的高层描述,实现混合构建-检索机制。在主流NaVQA数据集上的评估显示,EmbodiedLGR-Agent在推理和查询时间上达到当前最优表现,全局任务准确率也保持竞争力。此外,该系统已成功部署于实体机器人,在真实人机交互中验证了其实用性,且视觉语言模型与构建-检索流水线均在本地运行。

原文摘要 · Abstract (English)

As the world of agentic artificial intelligence applied to robotics evolves, the need for agents capable of building and retrieving memories and observations efficiently is increasing. Robots operating in complex environments must build memory structures to enable useful human-robot interactions by leveraging the mnemonic representation of the current operating context. People interacting with robots may expect the embodied agent to provide information about locations, events, or objects, which requires the agent to provide precise answers within human-like inference times to be perceived as responsive. We propose the Embodied Light Graph Retrieval Agent (EmbodiedLGR-Agent), a visual-language model (VLM)-driven agent architecture that constructs dense and efficient representations of robot operating environments. EmbodiedLGR-Agent directly addresses the need for an efficient memory representation of the environment by providing a hybrid building-retrieval approach built on parameter-efficient VLMs that store low-level information about objects and their positions in a semantic graph, while retaining high-level descriptions of the observed scenes with a traditional retrieval-augmented architecture. EmbodiedLGR-Agent is evaluated on the popular NaVQA dataset, achieving state-of-the-art performance in inference and querying times for embodied agents, while retaining competitive accuracy on the global task relative to the current state-of-the-art approaches. Moreover, EmbodiedLGR-Agent was successfully deployed on a physical robot, showing practical utility in real-world contexts through human-robot interaction, while running the visual-language model and the building-retrieval pipeline locally.

机器人记忆存储视觉语言模型实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。