arXiv:2504.03524cs.CVcs.IR2025-04被引 4

让机器人通过查询历史数据实现跨任务零样本导航

RANa: Retrieval-Augmented Navigation

  • 用检索增强机制让机器人调用过往经验
  • 在多个导航任务上实现零样本迁移并提升性能
  • 适合需要长期记忆的智能机器人应用

基于大规模学习的导航方法通常将每个任务视为新问题,代理在未知环境中从零开始。但在真实场景中,机器人应能利用此前操作积累的信息。为此,我们提出一种基于强化学习的检索增强型代理,可查询同一环境中的历史数据,并学会融合这些上下文信息。该方法采用全新架构,在ImageNav、Instance-ImageNav和ObjectNav上进行评估。检索与上下文编码均基于数据驱动,使用视觉基础模型(FM)实现语义与几何理解。我们构建了新基准,证明检索可实现跨任务与环境的零样本迁移,并显著提升表现。

原文摘要 · Abstract (English)

Methods for navigation based on large-scale learning typically treat each episode as a new problem, where the agent is spawned with a clean memory in an unknown environment. While these generalization capabilities to an unknown environment are extremely important, we claim that, in a realistic setting, an agent should have the capacity of exploiting information collected during earlier robot operations. We address this by introducing a new retrieval-augmented agent, trained with RL, capable of querying a database collected from previous episodes in the same environment and learning how to integrate this additional context information. We introduce a unique agent architecture for the general navigation task, evaluated on ImageNav, Instance-ImageNav and ObjectNav. Our retrieval and context encoding methods are data-driven and employ vision foundation models (FM) for both semantic and geometric understanding. We propose new benchmarks for these settings and we show that retrieval allows zero-shot transfer across tasks and environments while significantly improving performance.

机器人导航检索增强零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。