用异构掩码自回归模型加速机器人视频生成与策略评估
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
- 通过异构预训练融合多机器人、多任务数据,提升视频建模泛化性
- 生成视频视觉保真度更高,推理速度比之前快15倍,支持实时运行
- 适合需要高效仿真和合成数据的机器人学习研究者使用
我们提出异构掩码自回归(HMA)模型,用于建模动作-视频动态,以生成高质量数据并支持机器人学习的规模化评估。由于需在多样场景中保持计算效率以实现实时运行,构建交互式视频世界模型和策略面临挑战。HMA利用来自不同机器人形态、领域和任务的观测与动作序列进行异构预训练,并采用掩码自回归机制生成量化或软标记的视频预测。实验表明,该模型在真实世界中比现有机器人视频生成模型具有更优的视觉保真度和可控性,且推理速度提升15倍。经过后训练,该模型可作为从低层级动作输入出发的视频仿真器,用于策略评估和合成数据生成。
原文摘要 · Abstract (English)
We propose Heterogeneous Masked Autoregression (HMA) for modeling action-video dynamics to generate high-quality data and evaluation in scaling robot learning. Building interactive video world models and policies for robotics is difficult due to the challenge of handling diverse settings while maintaining computational efficiency to run in real time. HMA uses heterogeneous pre-training from observations and action sequences across different robotic embodiments, domains, and tasks. HMA uses masked autoregression to generate quantized or soft tokens for video predictions. \ourshort achieves better visual fidelity and controllability than the previous robotic video generation models with 15 times faster speed in the real world. After post-training, this model can be used as a video simulator from low-level action inputs for evaluating policies and generating synthetic data. See this link https://liruiw.github.io/hma for more information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。