arXiv:2601.17507cs.RO2026-01

用分层世界模型让机器人高效理解并执行高阶指令。

MetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level Instructions

  • 分层设计:语义层用VLM,物理层用隐空间动态模型。
  • 实验显示任务完成率与动作连贯性优于现有世界模型强化学习方法。
  • 适合研究具身智能、人形机器人指令执行的学者参考。

人形机器人在操作与移动任务中仍受语义-物理鸿沟制约。当前方法存在三大局限:强化学习样本效率低、模仿学习泛化能力差、视觉语言模型(VLM)生成动作物理不一致。本文提出MetaWorld,一种分层世界模型,通过专家策略迁移实现语义规划与物理控制的融合。该框架将任务解耦为由VLM驱动的语义层和在紧凑状态空间运行的隐动态模型。通过动态专家选择与运动先验融合机制,利用预训练多专家策略库作为可迁移知识,支持基于两阶段框架的高效在线适配。VLM作为语义接口,将指令映射为可执行技能,绕过符号接地问题。在Humanoid-Bench上的实验表明,MetaWorld在任务完成率与动作连贯性方面均优于基于世界模型的强化学习方法。

原文摘要 · Abstract (English)

Humanoid robot loco-manipulation remains constrained by the semantic-physical gap. Current methods face three limitations: Low sample efficiency in reinforcement learning, poor generalization in imitation learning, and physical inconsistency in VLMs. We propose MetaWorld, a hierarchical world model that integrates semantic planning and physical control via expert policy transfer. The framework decouples tasks into a VLM-driven semantic layer and a latent dynamics model operating in a compact state space. Our dynamic expert selection and motion prior fusion mechanism leverages a pre-trained multi-expert policy library as transferable knowledge, enabling efficient online adaptation via a two-stage framework. VLMs serve as semantic interfaces, mapping instructions to executable skills and bypassing symbol grounding. Experiments on Humanoid-Bench show MetaWorld outperforms world model-based RL in task completion and motion coherence. Our code will be found at https://anonymous.4open.science/r/metaworld-2BF4/

人形机器人分层模型指令执行视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。