让机器人通过理解任务进度生成下一步视觉目标,提升操作鲁棒性。
Incorporating Task Progress Knowledge for Subgoal Generation in Robotic Manipulation through Image Edits
- 结合任务进度信息,用图像编辑生成下一步视觉子目标。
- 在CALVIN基准上达到当前最优性能,适应不同初始姿态和速度。
- 适合研究具身智能、机器人规划与视觉导航的开发者。
理解任务进展使人类不仅能追踪已完成步骤,还能更好规划未来目标。我们提出TaKSIE框架,将任务进度知识融入机器人操作中的视觉子目标生成。通过联合训练循环网络与潜在扩散模型,根据当前观测和语言指令生成下一视觉子目标。执行时,机器人利用视觉进度表示监控任务进展,并自适应地从模型中采样下一视觉子目标,以指导操作策略。我们在仿真与真实机器人任务中训练并验证该模型,在CALVIN操作基准上取得当前最优表现。结果表明,引入任务进度知识可显著提升策略对不同初始机器人姿态或演示运动速度的鲁棒性。
原文摘要 · Abstract (English)
Understanding the progress of a task allows humans to not only track what has been done but also to better plan for future goals. We demonstrate TaKSIE, a novel framework that incorporates task progress knowledge into visual subgoal generation for robotic manipulation tasks. We jointly train a recurrent network with a latent diffusion model to generate the next visual subgoal based on the robot's current observation and the input language command. At execution time, the robot leverages a visual progress representation to monitor the task progress and adaptively samples the next visual subgoal from the model to guide the manipulation policy. We train and validate our model in simulated and real-world robotic tasks, achieving state-of-the-art performance on the CALVIN manipulation benchmark. We find that the inclusion of task progress knowledge can improve the robustness of trained policy for different initial robot poses or various movement speeds during demonstrations. The project website can be found at https://live-robotics-uva.github.io/TaKSIE/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。