用大模型生成任务奖励,让机器人更懂长时序操作
RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation
- 用大模型解析指令生成密集奖励,拆解长任务为子步骤
- 结合多视角自编码器,自动识别关键动作帧提升感知能力
- 在两个基准上显著提升成功率,尤其对长序列任务效果突出
长时序机器人操作的高效控制面临复杂表征与策略学习的挑战。基于模型的视觉强化学习虽具潜力,但在稀疏奖励和复杂视觉特征环境下仍表现受限。为此,本文提出针对长时序任务的RSPA流程,并构建了面向长时序机器人操作的LLM辅助多视角世界模型RoboHorizon。在RoboHorizon中,预训练大模型根据任务语言指令生成多阶段子任务的密集奖励结构,帮助机器人更好识别长时序任务;同时将关键帧发现融入多视角掩码自编码器(MAE)架构,增强对关键任务序列的感知能力。基于这些密集奖励与多视角表示,构建机器人世界模型,实现高效长时序规划,并通过强化学习算法可靠执行。在RLBench和FurnitureBench两个代表性基准上的实验表明,RoboHorizon优于现有先进视觉模型基强化学习方法,在RLBench的4个短时序任务上成功率达23.35%提升,在6个长时序任务及FurnitureBench的3个家具组装任务上提升达29.23%。
原文摘要 · Abstract (English)
Efficient control in long-horizon robotic manipulation is challenging due to complex representation and policy learning requirements. Model-based visual reinforcement learning (RL) has shown great potential in addressing these challenges but still faces notable limitations, particularly in handling sparse rewards and complex visual features in long-horizon environments. To address these limitations, we propose the Recognize-Sense-Plan-Act (RSPA) pipeline for long-horizon tasks and further introduce RoboHorizon, an LLM-assisted multi-view world model tailored for long-horizon robotic manipulation. In RoboHorizon, pre-trained LLMs generate dense reward structures for multi-stage sub-tasks based on task language instructions, enabling robots to better recognize long-horizon tasks. Keyframe discovery is then integrated into the multi-view masked autoencoder (MAE) architecture to enhance the robot's ability to sense critical task sequences, strengthening its multi-stage perception of long-horizon processes. Leveraging these dense rewards and multi-view representations, a robotic world model is constructed to efficiently plan long-horizon tasks, enabling the robot to reliably act through RL algorithms. Experiments on two representative benchmarks, RLBench and FurnitureBench, show that RoboHorizon outperforms state-of-the-art visual model-based RL methods, achieving a 23.35% improvement in task success rates on RLBench's 4 short-horizon tasks and a 29.23% improvement on 6 long-horizon tasks from RLBench and 3 furniture assembly tasks from FurnitureBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。