arXiv:2510.24118cs.ROcs.AI2025-10被引 1

用语言3D高斯点云构建记忆,让机器人能听懂多模态指令导航到多个目标。

LagMemo: Language 3D Gaussian Splatting Memory for Multi-modal Open-vocabulary Multi-goal Visual Navigation

  • 用语言3D高斯点云构建统一空间语义记忆,支持跨模态理解。
  • 在多目标导航任务中准确率显著优于现有方法,提升明显。
  • 适合需要理解自然语言指令的智能机器人导航研究者使用。

利用视觉信息导航至指定目标是智能机器人的一项基础能力。为应对多模态、开放词汇目标查询及多目标视觉导航的实际需求,我们提出LagMemo,一种基于语言3D高斯点云记忆的导航系统。在一次探索中,LagMemo构建具有强空间-语义关联的统一3D语言记忆。面对新任务目标时,系统可高效查询记忆,预测候选目标位置,并通过局部感知验证机制动态匹配与确认目标。为实现公平严谨评估,我们从GOAT-Bench中提炼出高质量的核心数据集GOAT-Core。实验表明,LagMemo的记忆模块有效实现了多模态开放词汇定位,在多目标视觉导航任务中显著优于当前最先进方法。

原文摘要 · Abstract (English)

Navigating to a designated goal using visual information is a fundamental capability for intelligent robots. To address the practical demands of multi-modal, open-vocabulary goal queries and multi-goal visual navigation, we propose LagMemo, a navigation system that leverages a language 3D Gaussian Splatting memory. During a one-time exploration, LagMemo constructs a unified 3D language memory with robust spatial-semantic correlations. With incoming task goals, the system efficiently queries the memory, predicts candidate goal locations, and integrates a local perception-based verification mechanism to dynamically match and validate goals. For fair and rigorous evaluation, we curate GOAT-Core, a high-quality core split distilled from GOAT-Bench. Experimental results show that LagMemo's memory module enables effective multi-modal open-vocabulary localization, and significantly outperforms state-of-the-art methods in multi-goal visual navigation. Project page: https://weekgoodday.github.io/lagmemo

视觉导航语言记忆3D高斯多目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。