通过人机双向模仿预训练,提升视觉-语言-动作模型的泛化能力
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
- 利用人体与机械臂动作空间的对齐机制实现双向行为模仿
- 在仿真和真实机器人上分别提升25%和14%的性能
- 适合需要跨平台泛化的机器人控制任务
尽管利用大量人类视频和模拟机器人数据可缓解真实机器人数据稀缺问题,现有视觉-语言-动作模型(VLAs)仍受限于摄像头视角、视觉外观及具身形态差异,导致泛化能力不足。为此,我们提出MiVLA,一种基于人机双向模仿预训练的通用型VLA,利用人体手部与机械臂的动作相似性,建立对人类行为与机器人控制的强行为先验。具体地,通过左右手坐标系的运动学规则,在人类与机器人动作空间间实现双向对齐。给定人类或模拟机器人示范,MiVLA可预测一种具身的行为轨迹,并模仿另一种未见示范的具身行为。基于此双向模仿,模型融合了真实人类数据的行为保真度与模拟机器人数据的操作多样性,显著提升下游任务泛化能力。在三种机器人(ARX、PiPer、LocoMan)的仿真与真实平台上的实验表明,MiVLA在仿真中相比最优基线(如π₀、π₀.₅、H-RDT)提升25%,在真实机器人控制任务中提升14%。
原文摘要 · Abstract (English)
While leveraging abundant human videos and simulated robot data poses a scalable solution to the scarcity of real-world robot data, the generalization capability of existing vision-language-action models (VLAs) remains limited by mismatches in camera views, visual appearance, and embodiment morphologies. To overcome this limitation, we propose MiVLA, a generalizable VLA empowered by human-robot mutual imitation pre-training, which leverages inherent behavioral similarity between human hands and robotic arms to build a foundation of strong behavioral priors for both human actions and robotic control. Specifically, our method utilizes kinematic rules with left/right hand coordinate systems for bidirectional alignment between human and robot action spaces. Given human or simulated robot demonstrations, MiVLA is trained to forecast behavior trajectories for one embodiment, and imitate behaviors for another one unseen in the demonstration. Based on this mutual imitation, it integrates the behavioral fidelity of real-world human data with the manipulative diversity of simulated robot data into a unified model, thereby enhancing the generalization capability for downstream tasks. Extensive experiments conducted on both simulation and real-world platforms with three robots (ARX, PiPer and LocoMan), demonstrate that MiVLA achieves strong improved generalization capability, outperforming state-of-the-art VLAs (e.g., $\boldsymbolπ_{0}$, $\boldsymbolπ_{0.5}$ and H-RDT) by 25% in simulation, and 14% in real-world robot control tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。