arXiv:2508.17198cs.AI2025-08被引 13

让机器人像人一样构建空间认知地图,实现智能导航

From reactive to cognitive: brain-inspired spatial intelligence for embodied agents

  • 模仿大脑结构构建三类空间记忆:地标、路径、全局地图
  • 结合大模型实现零样本跨任务导航,效率与准确率均达顶尖水平
  • 适合研究具身智能、自主导航的学者与工程师

空间认知通过构建空间内部模型,支持自适应的目标导向行为。生物系统将空间知识整合为三种相互关联的形式:用于显著线索的地标、用于运动轨迹的路径知识,以及用于地图式表示的全景知识。尽管多模态大语言模型(MLLM)已在具身智能体中实现视觉-语言推理,但这些方法缺乏结构化空间记忆,仅能被动响应,限制了其在复杂真实环境中的泛化与适应能力。本文提出脑启发的空间认知导航框架(BSC-Nav),统一构建并利用具身智能体的结构化空间记忆。BSC-Nav从自身视角轨迹和上下文线索生成以客体为中心的认知地图,并动态检索与语义目标对齐的空间知识。集成强大MLLM后,BSC-Nav在多种导航任务中达到当前最优性能,展现强零样本泛化能力,支持现实世界中的多样化具身行为,为通用空间智能提供可扩展且生物合理的路径。

原文摘要 · Abstract (English)

Spatial cognition enables adaptive goal-directed behavior by constructing internal models of space. Robust biological systems consolidate spatial knowledge into three interconnected forms: \textit{landmarks} for salient cues, \textit{route knowledge} for movement trajectories, and \textit{survey knowledge} for map-like representations. While recent advances in multi-modal large language models (MLLMs) have enabled visual-language reasoning in embodied agents, these efforts lack structured spatial memory and instead operate reactively, limiting their generalization and adaptability in complex real-world environments. Here we present Brain-inspired Spatial Cognition for Navigation (BSC-Nav), a unified framework for constructing and leveraging structured spatial memory in embodied agents. BSC-Nav builds allocentric cognitive maps from egocentric trajectories and contextual cues, and dynamically retrieves spatial knowledge aligned with semantic goals. Integrated with powerful MLLMs, BSC-Nav achieves state-of-the-art efficacy and efficiency across diverse navigation tasks, demonstrates strong zero-shot generalization, and supports versatile embodied behaviors in the real physical world, offering a scalable and biologically grounded path toward general-purpose spatial intelligence.

具身智能空间认知导航大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。