让机器人在导航失败后自主进化计划,提升成功率与效率。
Agentic Self-Evolutionary Replanning for Embodied Navigation
- 用自演化动作模型实时学习经验,动态调整策略。
- 在多个基准测试中成功率更高,令牌消耗更低。
- 适合需要高鲁棒性的复杂环境导航任务。
复杂环境中机器人导航失败不可避免。为增强适应性,重规划(RP)允许机器人在失败后调整路径直至成功。现有方法固定动作模型,错失自我优化机会。本文提出自演化重规划(SERP),通过运行时学习实现模型动态进化。不同于依赖预设参数的静态进化方式,我们引入基于上下文学习与自动微分(ILAD)的智能体自演化动作模型,实现函数自适应调整与全局参数重置。为实现高效重规划,还提出基于大语言模型(LLM)推理的图链式思维(GCOT)方法,对压缩图进行推理。大量仿真与真实世界实验表明,SERP在多个基准上实现了更高的成功率与更低的令牌开销,验证了其在多样化环境中的卓越鲁棒性与效率。
原文摘要 · Abstract (English)
Failure is inevitable for embodied navigation in complex environments. To enhance the resilience, replanning (RP) is a viable option, where the robot is allowed to fail, but is capable of adjusting plan until success. However, existing RP approaches freeze the ego action model and miss the opportunities to explore better plans by upgrading the robot itself. To address this limitation, we propose Self-Evolutionary RePlanning, or SERP for short, which leads to a paradigm shift from frozen models towards evolving models by run-time learning from recent experiences. In contrast to existing model evolution approaches that often get stuck at predefined static parameters, we introduce agentic self-evolving action model that uses in-context learning with auto-differentiation (ILAD) for adaptive function adjustment and global parameter reset. To achieve token-efficient replanning for SERP, we also propose graph chain-of-thought (GCOT) replanning with large language model (LLM) inference over distilled graphs. Extensive simulation and real-world experiments demonstrate that SERP achieves higher success rate with lower token expenditure over various benchmarks, validating its superior robustness and efficiency across diverse environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。