构建人形机器人世界模型,实现未来视觉与状态的精准预测。
Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report
- 用视频生成模型结合机器人状态进行未来帧预测。
- 采样任务达23.0 dB PSNR,压缩任务Top-500 CE为6.6386。
- 首个开源真实人形交互基准,适合机器人感知与生成研究者。
世界模型是人工智能与机器人领域的重要范式,使智能体能够通过预测视觉观测或紧凑的潜在状态来推理未来。1X世界模型挑战赛推出了一个开源的真实人形机器人交互基准,包含两个互补赛道:采样(聚焦未来图像帧预测)和压缩(聚焦未来离散潜在码预测)。在采样赛道,我们采用视频生成基础模型Wan-2.2 TI2V-5B,通过AdaLN-Zero将视频生成条件于机器人状态,并使用LoRA进行后续微调。在压缩赛道,我们从零训练了一个时空变换器模型。我们的模型在采样任务中达到23.0 dB PSNR,压缩任务中取得Top-500 CE 6.6386,两项均获第一名。
原文摘要 · Abstract (English)
World models are a powerful paradigm in AI and robotics, enabling agents to reason about the future by predicting visual observations or compact latent states. The 1X World Model Challenge introduces an open-source benchmark of real-world humanoid interaction, with two complementary tracks: sampling, focused on forecasting future image frames, and compression, focused on predicting future discrete latent codes. For the sampling track, we adapt the video generation foundation model Wan-2.2 TI2V-5B to video-state-conditioned future frame prediction. We condition the video generation on robot states using AdaLN-Zero, and further post-train the model using LoRA. For the compression track, we train a Spatio-Temporal Transformer model from scratch. Our models achieve 23.0 dB PSNR in the sampling task and a Top-500 CE of 6.6386 in the compression task, securing 1st place in both challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。