让通用机械臂策略在无人机上精准执行,无需重新训练。
UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies
- 用人类手持抓取数据训练通用视觉运动策略,部署时动态适配不同机器人形态。
- 在无人机任务中成功率提升37%,抗干扰能力显著增强。
- 适合希望快速部署通用操控技能到受限机器人平台的研究者与工程师。
我们提出UMI-on-Air框架,实现对无体感(embodiment-agnostic)操纵策略的体感感知部署。该方法利用多种非受限人类示范(通过手持夹爪收集,称为UMI)训练可泛化的视觉运动策略。将这些策略迁移到受约束机器人形态(如空中机械臂)时,常因控制与动力学不匹配导致分布外行为和执行失败。为此,我们提出体感感知扩散策略(EADP),在推理时将高层的UMI策略与低层体感特异性控制器耦合。通过将控制器追踪成本的梯度反馈至扩散采样过程,引导轨迹生成向适应部署体感的动力学可行模式。该方法实现测试时即插即用的体感感知轨迹自适应。我们在多个长周期、高精度空中操作任务中验证了该方法,相比无引导扩散基线,成功率达37%提升,效率与鲁棒性显著改善。最后,我们在此前未见过的环境中部署,使用野外采集的UMI示范,展示了跨多样甚至高度受限体感规模化通用操控技能的实用路径。所有代码、数据、检查点及结果视频见 umi-on-air.github.io。
原文摘要 · Abstract (English)
We introduce UMI-on-Air, a framework for embodiment-aware deployment of embodiment-agnostic manipulation policies. Our approach leverages diverse, unconstrained human demonstrations collected with a handheld gripper (UMI) to train generalizable visuomotor policies. A central challenge in transferring these policies to constrained robotic embodiments-such as aerial manipulators-is the mismatch in control and robot dynamics, which often leads to out-of-distribution behaviors and poor execution. To address this, we propose Embodiment-Aware Diffusion Policy (EADP), which couples a high-level UMI policy with a low-level embodiment-specific controller at inference time. By integrating gradient feedback from the controller's tracking cost into the diffusion sampling process, our method steers trajectory generation towards dynamically feasible modes tailored to the deployment embodiment. This enables plug-and-play, embodiment-aware trajectory adaptation at test time. We validate our approach on multiple long-horizon and high-precision aerial manipulation tasks, showing improved success rates, efficiency, and robustness under disturbances compared to unguided diffusion baselines. Finally, we demonstrate deployment in previously unseen environments, using UMI demonstrations collected in the wild, highlighting a practical pathway for scaling generalizable manipulation skills across diverse-and even highly constrained-embodiments. All code, data, checkpoints, and result videos can be found at umi-on-air.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。