让机器人在不同形态下通用,通过设计等变动作空间提升泛化能力
Toward Embodiment Equivariant Vision-Language-Action Policy
- 构建动作空间与策略的形态等变理论,使模型对机器人结构变化保持不变性
- 提出等变动作解码器,实现跨形态配置的统一控制策略
- 适合需要快速适配新机器人硬件的研究者和工业应用
视觉-语言-动作策略通过大规模预训练学习跨任务、环境和机器人形态的操控技能。然而,其在新机器人配置下的泛化能力仍受限。现有方法多关注模型规模、数据集规模与多样性,却忽视动作空间设计,导致配置泛化问题,需高昂代价进行适应。本文将跨形态预训练建模为对形态变换保持等变的策略设计,提出三项关键改进:(i) 建立动作空间与策略设计的形态等变理论;(ii) 设计等变动作解码器以强制配置等变性;(iii) 引入几何感知网络架构,增强对形态无关的空间推理能力。大量仿真与真实世界实验表明,该方法显著提升预训练效果,并实现对新型机器人形态的高效微调。代码已公开于 https://github.com/hhcaz/e2vla。
原文摘要 · Abstract (English)
Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches emphasize model size, dataset scale and diversity while paying less attention to the design of action spaces. This leads to the configuration generalization problem, which requires costly adaptation. We address this challenge by formulating cross-embodiment pre-training as designing policies equivariant to embodiment configuration transformations. Building on this principle, we propose a framework that (i) establishes a embodiment equivariance theory for action space and policy design, (ii) introduces an action decoder that enforces configuration equivariance, and (iii) incorporates a geometry-aware network architecture to enhance embodiment-agnostic spatial reasoning. Extensive experiments in both simulation and real-world settings demonstrate that our approach improves pre-training effectiveness and enables efficient fine-tuning on novel robot embodiments. Our code is available at https://github.com/hhcaz/e2vla
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。