用自然语言指令训练人形机器人行走,实现稳定高效运动。
WOLF-VLA: Whole-Body Humanoid Optimal Locomotion Framework for Vision-Language-Action Learning

- 结合最优控制与多模态数据,生成动态一致的运动轨迹。
- 在六类任务中表现优异,对初始条件变化有强鲁棒性。
- 适合研究指令驱动的人形机器人运动迁移与泛化。
视觉-语言-动作(VLA)模型在机器人操作中展现出强大泛化能力,但其在全身、接触密集型人形机器人行走中的应用仍因数据稀缺、缺乏动态一致性示范,以及难以在学习框架中编码最优性与安全性而受到严重限制。本文提出统一框架 WOLF-VLA,将全身最优控制(OC)运动生成与大规模多模态数据集相结合,训练出可直接从自然语言指令生成人形机器人行走策略的 VLA 模型。我们构建了一个涵盖六类行走相关任务的完整数据集,每类任务参数化于环境变化、物体颜色、位置及视觉干扰因素。通过联合使用采集的关节轨迹、第一视角视觉观测与自然语言指令训练该模型,所得策略在多种任务和环境设置下表现出强推理能力与对初始条件变化的鲁棒性,并达到竞争性性能。系统性消融实验验证了各模态对模型表现的影响。完整数据集、模型检查点与基准仿真套件将开源发布,建立可复现的动态一致基准,推动全身人形机器人视觉-语言-动作控制的可扩展迁移研究。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently demonstrated strong generalization in robotic manipulation, yet their applicability to whole-body, contact-rich humanoid locomotion remains severely underexplored due to data scarcity, the absence of dynamically consistent demonstrations, and the difficulty of encoding optimality and safety in learning-based pipelines. This work introduces a unified framework WOLF-VLA that integrates whole-body optimal-control (OC) motion synthesis with large-scale multi-modal dataset to train VLAs capable of generating humanoid locomotion policies directly from natural-language instructions. We construct a comprehensive dataset of dynamically feasible humanoid trajectories across six locomotion-related task families, each parameterized by environmental variations, object colors, placements, and visual distractors. We train a VLA model using the collected joint trajectories, ego-centric visual observations and natural language instruction, yielding a policy that exhibits strong reasoning and robustness to initial-condition variability, and competitive performance across several tasks and environment settings. A systematic ablation study demonstrates the impact of each modality on the model performance. The full dataset, model checkpoints, and benchmarking simulation suite will be openly released, establishing a reproducible dynamically consistent benchmark for whole-body humanoid locomotion rich VLA control and enabling future research in scalable transfer of instruction-driven locomotion policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。