RePlan-Bot通过多层级重规划提升机器人执行长任务的可靠性。
RePlan-Bot: Multi-Level Replanning for Embodied Instruction Following

- 引入分层重规划机制,动态调整子目标并实时修正动作。
- 在ALFRED基准上实现最优成功率,未见环境表现提升12.3%。
- 适合需要高鲁棒性的具身智能任务,如家庭服务机器人。
具身指令遵循(EIF)要求智能体在交互式3D环境中理解并执行复杂的自然语言指令。尽管近期取得进展,现有方法在长时程规划和处理不可逆状态变化方面仍表现不佳,导致任务成功率较低。为此,我们提出RePlan-Bot,一种新型的EIF智能体,可在任务执行过程中进行多层级、持续的重规划。RePlan-Bot融合了基于LLM的高层级审计器,根据环境反馈动态调整子目标;基于多层实例地图的常识引导搜索机制,实现精确且结构化的物体定位;以及轻量级ViT-based校正器,提前修复高风险底层动作。在ALFRED基准上的评估显示,RePlan-Bot在已见与未见环境中均达到当前最佳性能,展现出卓越的适应性与可靠性。
原文摘要 · Abstract (English)
Embodied instruction following (EIF) requires agents to understand and execute complex natural language commands within interactive 3D environments. Despite recent advances, existing methods often fail in long-horizon planning and handling irreversible state changes, resulting in low task success rates. To address these challenges, we introduce RePlan-Bot, a novel EIF agent that performs multi-level, continuous replanning throughout task execution. RePlan-Bot integrates a high-level LLM-based auditor for dynamic sub-goal adjustments guided by environmental feedback, a commonsense-guided search mechanism based on a multi-layered instance map for precise and structured object localization, and a lightweight ViT-based corrector to preemptively fix risky low-level actions. Evaluated on the ALFRED benchmark, RePlan-Bot achieves state-of-the-art performance in both seen and unseen environments, demonstrating superior adaptability and reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。