arXiv:2603.00182cs.ROcs.AI2026-03被引 2

让Transformer模型理解机器人身体结构,提升跨机型泛化能力。

Embedding Morphology into Transformers for Cross-Robot Policy Learning

  • 用关节动作令牌和分块时间压缩,显式建模身体结构
  • 通过拓扑注意力偏置,引导信息沿机械连接路径传播
  • 结合关节属性描述,捕捉超越连接关系的语义信息

跨机器人策略学习——训练一个策略在多种机器人形态上均表现良好——仍是机器人学习中的核心挑战。基于Transformer的策略(如视觉-语言-动作模型)通常不依赖具体形态,需仅从观测中推断运动学结构,这会降低不同形态间的鲁棒性,甚至影响单个形态内的性能。本文提出一种形态感知的Transformer策略,通过三种机制注入形态信息:(1) 运动学令牌,将动作按关节分解,并通过每关节的时间分块压缩序列长度;(2) 拓扑感知注意力偏置,在自注意力中引入运动学拓扑作为归纳偏置,鼓励信息沿机械连接边传递;(3) 关节属性条件,用每个关节的描述符增强拓扑信息,捕捉连接关系之外的语义。在多种机器人形态上,该结构化集成方法持续优于基础的pi0.5 VLA模型,表明其在单个形态内及跨形态间均具有更强鲁棒性。

原文摘要 · Abstract (English)

Cross-robot policy learning -- training a single policy to perform well across multiple embodiments -- remains a central challenge in robot learning. Transformer-based policies, such as vision-language-action (VLA) models, are typically embodiment-agnostic and must infer kinematic structure purely from observations, which can reduce robustness across embodiments and even limit performance within a single embodiment. We propose an embodiment-aware transformer policy that injects morphology via three mechanisms: (1) kinematic tokens that factorize actions across joints and compress time through per-joint temporal chunking; (2) a topology-aware attention bias that encodes kinematic topology as an inductive bias in self-attention, encouraging message passing along kinematic edges; and (3) joint-attribute conditioning that augments topology with per-joint descriptors to capture semantics beyond connectivity. Across a range of embodiments, this structured integration consistently improves performance over a vanilla pi0.5 VLA baseline, indicating improved robustness both within an embodiment and across embodiments.

机器人学习Transformer形态感知策略泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。