用动作指令连接语义反思与具体操作,让机器人更精准地自我修正。
Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action Correction
- 通过动作指令作为桥梁,将高层语义反思转化为可执行的动作调整。
- 结合多任务扩散策略,实现基于视觉反馈的高频精细动作修正。
- 适合需要长期自适应、高精度操作的机器人系统研发人员。
构建通用的自我纠错系统对机器人应对失败至关重要。尽管多模态大语言模型(MLLMs)赋予了机器人语义层面的反思能力,但如何将这种反思转化为具体、细粒度的机器人动作修正仍是重大挑战。为此,我们提出了Phoenix框架,利用动作指令作为桥梁,连接高层语义反思与底层动作修正。该框架采用双过程动作调整机制,由MLLMs驱动,将语义反思转化为粗粒度动作指令调整;为进一步指导细粒度动作修正,提出一种多任务运动条件扩散策略,融合视觉观测实现高频修正。通过两模型协同,将泛化能力需求从底层操作策略转移到由MLLMs驱动的动作调整模型,从而实现精确的细粒度动作修正。此外,还设计了一种终身学习方法,使模型能通过与动态环境的交互持续提升性能。在RoboMimic仿真及真实场景中的实验验证了该框架在多种操作任务中具备优异的泛化性与鲁棒性。
原文摘要 · Abstract (English)
Building a generalizable self-correction system is crucial for robots to recover from failures. Despite advancements in Multimodal Large Language Models (MLLMs) that empower robots with semantic reflection ability for failure, translating semantic reflection into how to correct fine-grained robotic actions remains a significant challenge. To address this gap, we build the Phoenix framework, which leverages motion instruction as a bridge to connect high-level semantic reflection with low-level robotic action correction. In this motion-based self-reflection framework, we start with a dual-process motion adjustment mechanism with MLLMs to translate the semantic reflection into coarse-grained motion instruction adjustment. To leverage this motion instruction for guiding how to correct fine-grained robotic actions, a multi-task motion-conditioned diffusion policy is proposed to integrate visual observations for high-frequency robotic action correction. By combining these two models, we could shift the demand for generalization capability from the low-level manipulation policy to the MLLMs-driven motion adjustment model and facilitate precise, fine-grained robotic action correction. Utilizing this framework, we further develop a lifelong learning method to automatically improve the model's capability from interactions with dynamic environments. The experiments conducted in both the RoboMimic simulation and real-world scenarios prove the superior generalization and robustness of our framework across a variety of manipulation tasks. Our code is released at \href{https://github.com/GeWu-Lab/Motion-based-Self-Reflection-Framework}{https://github.com/GeWu-Lab/Motion-based-Self-Reflection-Framework}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。