arXiv:2504.15643cs.RO2025-04综述被引 2

综述多模态感知在目标导向导航中的应用与挑战

Multimodal Perception for Goal-oriented Navigation: A Survey

  • 按推理域统一分析视觉、语言、声音信息的使用方式
  • 梳理200篇论文,揭示不同导航方法的共性与差异
  • 适合研究智能机器人导航与多模态融合的学者

目标导向导航是自主系统面临的基础挑战,要求智能体在复杂环境中导航至指定目标。本文从推理域的统一视角出发,全面分析多模态导航方法,探讨智能体如何利用视觉、语言和听觉信息进行环境感知、推理与导航。主要贡献包括:基于推理机制对导航方法进行分类;系统分析共享计算基础如何支撑不同任务的多样化方法;识别各类导航范式中的共性模式与独特优势;并探讨多模态感知集成的挑战与机遇。此外,本文回顾了约200篇相关文献,深入剖析当前研究格局。

原文摘要 · Abstract (English)

Goal-oriented navigation presents a fundamental challenge for autonomous systems, requiring agents to navigate complex environments to reach designated targets. This survey offers a comprehensive analysis of multimodal navigation approaches through the unifying perspective of inference domains, exploring how agents perceive, reason about, and navigate environments using visual, linguistic, and acoustic information. Our key contributions include organizing navigation methods based on their primary environmental reasoning mechanisms across inference domains; systematically analyzing how shared computational foundations support seemingly disparate approaches across different navigation tasks; identifying recurring patterns and distinctive strengths across various navigation paradigms; and examining the integration challenges and opportunities of multimodal perception to enhance navigation capabilities. In addition, we review approximately 200 relevant articles to provide an in-depth understanding of the current landscape.

导航多模态智能体综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。