用物体运动轨迹统一机器人动作与控制,跨数据源预训练后可支持多种任务。
Unified Motion-Action Modeling for Heterogeneous Robot Learning

- 以物体3D运动轨迹为共享接口,联合建模机器人动作与运动变化。
- 在混合数据上预训练,少样本演示下实现任务迁移与动态建模。
- 无需标注任务指令,适合多场景异构数据的机器人学习应用。
我们提出统一运动-动作(UMA)模型,通过3D物体运动轨迹作为共享接口,连接视觉-运动控制与动力学建模。UMA将物体运动与机器人动作视为在掩码生成目标下共同演化的变量,掩码模式决定预训练监督方式与部署推理模式。利用事后重标注的运动上下文与对比学习目标,分离任务意图与场景几何,使UMA可在无手动标注任务指令的情况下,跨异构数据源进行多任务预训练。部署时,同一套预训练参数支持运动条件下的视觉-运动控制、基于运动的动力学建模以及少样本示范的任务适应。在机器人示范、人类视频和仿真数据混合数据上预训练,UMA在各类推理模式下均优于现有专用基线方法。
原文摘要 · Abstract (English)
We present Unified Motion-Action (UMA) Model, an approach that uses 3D object motion trajectories as a shared interface to bridge visuomotor control and dynamics modeling. UMA treats object motion and robot actions as co-evolving variables under a masked generative objective, in which the mask pattern determines both the supervision regime during pretraining and the inference mode at deployment. Using hindsight-relabeled motion contexts and a contrastive objective that disentangles task intent from scene geometry, UMA enables multi-task pretraining across heterogeneous data sources without requiring manually annotated task instructions. At deployment, the same pretrained parameters support motion-conditioned visuomotor control, motion-based dynamics modeling, and task adaptation from few-shot demonstrations. Pretrained on a mixture of robot demonstrations, human videos, and simulated data, UMA consistently outperforms state-of-the-art baselines specialized for each inference mode.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。