用生成式预期实现机器人闭环操控,提升长任务鲁棒性
Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation

- 用文本控制的视频扩散模型生成视觉计划作为参考
- 在真实机器人任务中提升8%性能,超越现有开环方法
- 适合需要实时反馈的复杂操作场景研究者
尽管近年来机器人与具身智能取得显著进展,但部署机器人完成长时序任务仍是重大挑战。多数前期工作遵循开环范式,缺乏实时反馈,导致误差累积和鲁棒性不足。少数方法尝试利用像素级差异或预训练视觉表征建立反馈机制,但其有效性与适应性受限。受经典闭环控制系统启发,我们提出CLOVER——一种融合反馈机制的闭环视觉-运动控制框架。该框架包含:文本条件视频扩散模型用于生成视觉计划作为参考输入、可度量的嵌入空间实现精准误差量化、以及基于反馈驱动的控制器,可根据反馈优化动作并按需启动重规划。CLOVER在真实机器人任务中表现显著提升,在CALVIN基准上较此前开环方法提升8%。代码与检查点已开源。
原文摘要 · Abstract (English)
Despite significant progress in robotics and embodied AI in recent years, deploying robots for long-horizon tasks remains a great challenge. Majority of prior arts adhere to an open-loop philosophy and lack real-time feedback, leading to error accumulation and undesirable robustness. A handful of approaches have endeavored to establish feedback mechanisms leveraging pixel-level differences or pre-trained visual representations, yet their efficacy and adaptability have been found to be constrained. Inspired by classic closed-loop control systems, we propose CLOVER, a closed-loop visuomotor control framework that incorporates feedback mechanisms to improve adaptive robotic control. CLOVER consists of a text-conditioned video diffusion model for generating visual plans as reference inputs, a measurable embedding space for accurate error quantification, and a feedback-driven controller that refines actions from feedback and initiates replans as needed. Our framework exhibits notable advancement in real-world robotic tasks and achieves state-of-the-art on CALVIN benchmark, improving by 8% over previous open-loop counterparts. Code and checkpoints are maintained at https://github.com/OpenDriveLab/CLOVER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。