用视频动作模型让机器人更懂物理,提升控制效率与泛化能力
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- 用预训练视频模型捕捉视觉动态和语义,减少对专家数据依赖
- 在模拟和真实任务中样本效率提升10倍,收敛速度加快2倍
- 适合需要高效学习物理交互的机器人控制研究者
当前主流的视觉-语言-动作模型(VLAs)基于大规模但静态的网络数据预训练,虽提升了语义泛化能力,却需从机器人轨迹中隐式推断复杂物理动态与时间依赖关系,导致持续依赖大量专家数据。我们指出,尽管视觉-语言预训练能捕捉语义先验,却忽视了物理因果性。为此,提出mimic-video:一种新型视频-动作模型(VAM),将互联网规模的视频模型与基于流匹配的动作解码器结合,后者以视频隐表示为条件生成低层机器人动作,充当逆动力学模型。实验表明,该方法在仿真和真实机器人操作任务中达到领先性能,相比传统VLA架构,样本效率提升10倍,收敛速度加快2倍。
原文摘要 · Abstract (English)
Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy must implicitly infer complex physical dynamics and temporal dependencies solely from robot trajectories. This reliance creates an unsustainable data burden, necessitating continuous, large-scale expert data collection to compensate for the lack of innate physical understanding. We contend that while vision-language pretraining effectively captures semantic priors, it remains blind to physical causality. A more effective paradigm leverages video to jointly capture semantics and visual dynamics during pretraining, thereby isolating the remaining task of low-level control. To this end, we introduce mimic-video, a novel Video-Action Model (VAM) that pairs a pretrained Internet-scale video model with a flow matching-based action decoder conditioned on its latent representations. The decoder serves as an Inverse Dynamics Model (IDM), generating low-level robot actions from the latent representation of video-space action plans. Our extensive evaluation shows that our approach achieves state-of-the-art performance on simulated and real-world robotic manipulation tasks, improving sample efficiency by 10x and convergence speed by 2x compared to traditional VLA architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。