arXiv:2604.16391cs.ROcs.CV2026-04被引 17

分离视觉前后向动力学预训练,提升机器人泛化能力

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

论文配图:Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining
图 1 · 摘自论文原文
  • 分步预训练前向与逆向动态模型,解耦视觉预测与动作推理
  • 在CALVIN和SimplerEnv上分别实现4.51平均任务长度和81.3%真实部署成功率
  • 适用于从无标签视频中学习通用机器人行为,适合大规模数据训练

视觉-语言-动作(VLA)模型在构建通用机器人方面展现出巨大潜力,但仍面临2D图像预测与3D动作推断不匹配的困境。此外,视觉与动作的耦合训练限制了模型从大规模、无动作标注的网络视频中学习。为此,我们提出DeFI框架,通过分离视觉前向与逆向动力学预训练,实现各自数据源的利用,使视频生成与动作预测解耦。我们引入通用前向动力学模型(GFDM),在多样人类与机器人视频上预训练以预测未来;以及通用逆动力学模型(GIDM),通过自监督学习从无标签视频转换中推断潜在动作。两者集成于统一架构,在下游任务上进行端到端微调。实验表明,该方法在CALVIN ABC-D和SimplerEnv上表现优异,实现平均任务长度4.51、SimplerEnv-Fractal基准51.2%成功率及真实场景81.3%成功率,显著优于现有方法。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have shown great potential in building generalist robots, but still face a dilemma-misalignment of 2D image forecasting and 3D action prediction. Besides, such a vision-action entangled training manner limits model learning from large-scale, action-free web video data. To address these issues, we propose DeFI, a novel framework that Decouples visual Forward and Inverse dynamics pretraining to exploit respective data sources, wherein video generation and action prediction are disentangled. We introduce the General Forward Dynamics Model (GFDM), pretrained on diverse human and robot videos for future prediction, and the General Inverse Dynamics Model (GIDM), trained via self-supervised learning to infer latent actions from unlabeled video transitions. These models are then integrated into a unified architecture for end-to-end finetuning on downstream tasks. In this manner, GFDM and GIDM first shine separately and then cooperate for mutual benefit. Extensive experiments on CALVIN ABC-D and SimplerEnv demonstrate state-of-the-art performance, with DeFI achieving an average task length of 4.51 for CALVIN, 51.2% success rate on SimplerEnv-Fractal benchmark and 81.3% success rate in real-world deployment, significantly outperforming prior methods.

机器人学习解耦表征自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。