arXiv:2506.01182cs.ROcs.AI2025-06被引 6

让机器人预测自身动作后的视觉变化,实现复杂环境下的自主决策。

Humanoid World Models: Open World Foundation Models for Humanoid Robotics

  • 用控制指令预测未来视角视频,构建轻量级世界模型。
  • 在100小时真人演示数据上训练,支持长时序规划与策略学习。
  • 仅需1-2张显卡即可训练部署,适合小型研究团队使用。

类人机器人因其人体形态,特别适合在人类设计的环境中交互。然而,使类人机器人在复杂开放世界中进行推理、规划与行动仍具挑战。世界模型能通过预测给定动作的未来结果,作为长期规划的动力学模型,并生成用于策略学习的合成数据。本文提出人类形世界模型(HWM),一系列轻量级、开源的模型,可基于类人机器人控制指令预测未来第一人称视频。我们在100小时类人机器人示范数据上训练了两种生成模型:掩码变换器(Masked Transformers)和流匹配(Flow-Matching)。同时探索不同注意力机制与参数共享策略的架构变体。参数共享技术使模型尺寸减少33%-53%,对性能与视觉保真度影响极小。HWM专为学术与小实验室场景设计,可在1-2张GPU上完成训练与部署。

原文摘要 · Abstract (English)

Humanoid robots, with their human-like form, are uniquely suited for interacting in environments built for people. However, enabling humanoids to reason, plan, and act in complex open-world settings remains a challenge. World models, models that predict the future outcome of a given action, can support these capabilities by serving as a dynamics model in long-horizon planning and generating synthetic data for policy learning. We introduce Humanoid World Models (HWM), a family of lightweight, open-source models that forecast future egocentric video conditioned on humanoid control tokens. We train two types of generative models, Masked Transformers and Flow-Matching, on 100 hours of humanoid demonstrations. Additionally, we explore architectural variants with different attention mechanisms and parameter-sharing strategies. Our parameter-sharing techniques reduce model size by 33-53% with minimal impact on performance or visual fidelity. HWMs are designed to be trained and deployed in practical academic and small-lab settings, such as 1-2 GPUs.

世界模型类人机器人生成模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。