arXiv:2503.08306cs.ROcs.CV2025-03CVPR被引 10

研究端到端训练机器人在真实环境中的视觉导航推理机制。

Reasoning in visual navigation of end-to-end trained agents: a dynamical systems approach

  • 用动态系统视角分析机器人在真实环境中的决策过程。
  • 发现模型具备有限视野内的精确规划能力,依赖隐式记忆。
  • 揭示价值函数与长期规划的关联,适合机器人与强化学习研究者。

端到端训练的智能体在逼真环境中实现了高阶推理和零样本或语言引导行为,但现有基准仍以仿真为主。本文聚焦快速移动的真实机器人,开展大规模实验,在真实环境中完成\numepisodes{}次导航任务,分析端到端训练产生的推理类型。重点研究了智能体在开环预测中学习到的现实动力学特性及其与感知的交互关系。分析其如何利用隐含记忆保持场景结构信息和探索过程中获取的数据。探测智能体的规划能力,发现其记忆中存在有限时域内较精确的规划证据。进一步通过后验分析表明,智能体学习到的价值函数与长期规划相关。综合来看,本研究揭示了计算机视觉与序列决策方法结合在机器人控制中带来的新能力。交互工具可访问:europe.naverlabs.com/research/publications/reasoning-in-visual-navigation-of-end-to-end-trained-agents。

原文摘要 · Abstract (English)

Progress in Embodied AI has made it possible for end-to-end-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or language-conditioned behavior, but benchmarks are still dominated by simulation. In this work, we focus on the fine-grained behavior of fast-moving real robots and present a large-scale experimental study involving \numepisodes{} navigation episodes in a real environment with a physical robot, where we analyze the type of reasoning emerging from end-to-end training. In particular, we study the presence of realistic dynamics which the agent learned for open-loop forecasting, and their interplay with sensing. We analyze the way the agent uses latent memory to hold elements of the scene structure and information gathered during exploration. We probe the planning capabilities of the agent, and find in its memory evidence for somewhat precise plans over a limited horizon. Furthermore, we show in a post-hoc analysis that the value function learned by the agent relates to long-term planning. Put together, our experiments paint a new picture on how using tools from computer vision and sequential decision making have led to new capabilities in robotics and control. An interactive tool is available at europe.naverlabs.com/research/publications/reasoning-in-visual-navigation-of-end-to-end-trained-agents.

视觉导航端到端机器人动态系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。