探究图像目标导航中真正关键的能力,发现导航训练可自发产生相对位姿估计。
What does really matter in image goal navigation?
- 通过强化学习端到端训练,让智能体自主学会视觉对比与方向判断。
- 实验表明导航性能与相对位姿估计能力存在显著正相关。
- 成果可迁移至更真实场景,对具身智能有重要启发意义。
图像目标导航需掌握两种技能:一是基础导航能力,包括自由空间检测、障碍物识别及基于内部表征的决策;二是通过对比观测图像与目标图像来计算方向信息。现有先进方法依赖专用图像匹配或计算机视觉模块的相对位姿预训练。本文通过大规模实验研究,检验了仅通过强化学习对完整智能体进行端到端训练是否足以解决该任务。结果表明,近期方法的成功部分受仿真环境设置影响,存在仿真捷径;但这些能力仍可在一定程度上迁移到更真实场景。此外,我们发现导航表现与探测到的(涌现的)相对位姿估计表现存在显著相关性,这是一项关键子技能。
原文摘要 · Abstract (English)
Image goal navigation requires two different skills: firstly, core navigation skills, including the detection of free space and obstacles, and taking decisions based on an internal representation; and secondly, computing directional information by comparing visual observations to the goal image. Current state-of-the-art methods either rely on dedicated image-matching, or pre-training of computer vision modules on relative pose estimation. In this paper, we study whether this task can be efficiently solved with end-to-end training of full agents with RL, as has been claimed by recent work. A positive answer would have impact beyond Embodied AI and allow training of relative pose estimation from reward for navigation alone. In this large experimental study we investigate the effect of architectural choices like late fusion, channel stacking, space-to-depth projections and cross-attention, and their role in the emergence of relative pose estimators from navigation training. We show that the success of recent methods is influenced up to a certain extent by simulator settings, leading to shortcuts in simulation. However, we also show that these capabilities can be transferred to more realistic setting, up to some extent. We also find evidence for correlations between navigation performance and probed (emerging) relative pose estimation performance, an important sub skill.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。