arXiv:2509.11865cs.ROcs.AI2025-09被引 3

轻量级扩散变换器实现多机器人跨形态灵巧操作,性能远超传统方法。

Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer

  • 用跨形态归一化将视觉、本体感知与语言融合到统一隐空间
  • 在相同算力下,分布内成功率88.95%,跨场景仍保持强鲁棒性
  • 适合追求高效通用机械臂控制的工业与研究场景

扩展Transformer策略和扩散模型已推动机器人操作发展,但在轻量化、跨形态学习场景中结合二者仍具挑战。本文研究影响扩散-变换器策略稳定性和性能的设计选择,提出Tenma——一种用于双臂控制的轻量级扩散-变换器。Tenma通过跨形态归一化将多视角RGB、本体感知与语言映射至共享隐空间;采用联合状态-时间编码器实现时序对齐观测学习并提升推理速度;设计优化的扩散动作解码器以增强训练稳定性和学习能力。在基准测试中,尽管使用中等规模数据,Tenma在相同算力下实现了88.95%的平均成功率达分布内性能,且在物体与场景变化下仍保持良好表现,显著优于基线(最佳仅为18.12%)。结果表明,多模态与跨形态学习策略极大提升了基于Transformer的模仿学习政策潜力。

原文摘要 · Abstract (English)

Scaling Transformer policies and diffusion models has advanced robotic manipulation, yet combining these techniques in lightweight, cross-embodiment learning settings remains challenging. We study design choices that most affect stability and performance for diffusion-transformer policies trained on heterogeneous, multimodal robot data, and introduce Tenma, a lightweight diffusion-transformer for bi-manual arm control. Tenma integrates multiview RGB, proprioception, and language via a cross-embodiment normalizer that maps disparate state/action spaces into a shared latent space; a Joint State-Time encoder for temporally aligned observation learning with inference speed boosts; and a diffusion action decoder optimized for training stability and learning capacity. Across benchmarks and under matched compute, Tenma achieves an average success rate of 88.95% in-distribution and maintains strong performance under object and scene shifts, substantially exceeding baseline policies whose best in-distribution average is 18.12%. Despite using moderate data scale, Tenma delivers robust manipulation and generalization, indicating the great potential for multimodal and cross-embodiment learning strategies for further augmenting the capacity of transformer-based imitation learning policies.

机器人操作扩散模型跨形态多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。