arXiv:2508.15354cs.RO2025-08综述被引 4

系统梳理具身导航中感知、社交与运动智能的最新进展

Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey

  • 提出TOFRA五阶段框架,整合感知到决策全流程
  • 对比主流平台与评估指标,揭示现有方法短板
  • 适合机器人、AI导航方向研究者参考

具身导航(EN)通过感知、社交与运动智能,使机器人在复杂环境中执行以自身为中心的任务。与依赖显式定位和预设地图的传统方法不同,EN利用第一人称视觉和类人交互策略。本文提出一个五阶段综合框架——过渡(Transition)、观察(Observation)、融合(Fusion)、奖励-策略构建(Reward-policy construction)与动作(Action),简称TOFRA。该框架系统梳理了当前前沿工作,对相关平台与评估指标进行批判性分析,并指出现有关键开放问题。相关研究列表可访问 https://github.com/Franky-X/Awesome-Embodied-Navigation。

原文摘要 · Abstract (English)

Embodied navigation (EN) advances traditional navigation by enabling robots to perform complex egocentric tasks through sensing, social, and motion intelligence. In contrast to classic methodologies that rely on explicit localization and pre-defined maps, EN leverages egocentric perception and human-like interaction strategies. This survey introduces a comprehensive EN formulation structured into five stages: Transition, Observation, Fusion, Reward-policy construction, and Action (TOFRA). The TOFRA framework serves to synthesize the current state of the art, provide a critical review of relevant platforms and evaluation metrics, and identify critical open research challenges. A list of studies is available at https://github.com/Franky-X/Awesome-Embodied-Navigation.

具身导航智能体多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。