arXiv:2603.08572cs.ROcs.AI2026-03

用专家分工+视觉语言模型,让机器人走路同时干活更自然稳定。

MetaWorld-X: Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation

  • 分拆任务为多个专用专家,模仿人类动作学习
  • 通过语义路由实现多阶段任务自动组合执行
  • 适合需要复杂动作协调的仿人机器人研究

让仿人机器人同时完成行走与操作(loco-manipulation)仍面临重大挑战。现有强化学习方法通常依赖单一整体策略学习多种技能,易导致跨技能梯度干扰和高自由度系统中的运动冲突,生成行为常不自然、不稳定且泛化能力差。为此,我们提出MetaWorld-X,一种基于分治原则的层次化世界建模框架。该方法将复杂控制问题分解为一组专用专家策略(SEP),每个专家在人类运动先验约束下通过模仿-约束强化学习训练,引入生物力学一致性归纳偏置,确保生成动作自然且物理合理。在此基础上,我们进一步设计由视觉-语言模型(VLM)监督的智能路由机制(IRM),实现语义驱动的专家组合。VLM引导的路由器根据高层任务语义动态整合专家策略,支持组合泛化与多阶段任务的自适应执行。

原文摘要 · Abstract (English)

Learning natural, stable, and compositionally generalizable whole-body control policies for humanoid robots performing simultaneous locomotion and manipulation (loco-manipulation) remains a fundamental challenge in robotics. Existing reinforcement learning approaches typically rely on a single monolithic policy to acquire multiple skills, which often leads to cross-skill gradient interference and motion pattern conflicts in high-degree-of-freedom systems. As a result, generated behaviors frequently exhibit unnatural movements, limited stability, and poor generalization to complex task compositions. To address these limitations, we propose MetaWorld-X, a hierarchical world model framework for humanoid control. Guided by a divide-and-conquer principle, our method decomposes complex control problems into a set of specialized expert policies (Specialized Expert Policies, SEP). Each expert is trained under human motion priors through imitation-constrained reinforcement learning, introducing biomechanically consistent inductive biases that ensure natural and physically plausible motion generation. Building upon this foundation, we further develop an Intelligent Routing Mechanism (IRM) supervised by a Vision-Language Model (VLM), enabling semantic-driven expert composition. The VLM-guided router dynamically integrates expert policies according to high-level task semantics, facilitating compositional generalization and adaptive execution in multi-stage loco-manipulation tasks.

仿人机器人动作生成多任务控制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。