arXiv:2605.25044cs.RO2026-05被引 1

让机器人跨形态通用,用扩散模型统一动作生成。

X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models

论文配图:X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用扩散模型设计统一动作头,支持不同机器人的动作生成。
  • 在RoboCasa和Isaac Gym上分别提升15.3%和12.5%性能。
  • 适合做跨机器人形态的通用策略学习研究者。

从跨形态数据中学习通用策略仍是机器人领域的核心挑战。尽管视觉-语言-动作(VLA)模型在大规模多样数据上预训练,但通常需针对具体机器人进行微调才能达到良好表现,严重限制了泛化能力与跨形态知识迁移。为此,本文聚焦于共享机器人本体但末端执行器异构的跨形态场景,提出X-DiffVLA——一种基于扩散模型的统一跨形态动作头。该模型利用扩散模型的生成优势,捕捉跨形态数据中的多样性与潜在关联。具体地,引入无分类器引导的形态强制机制(Embodiment Forcing),隐式引导动作生成聚焦于特定形态的功能组件,无需显式标注即可捕获细微结构差异;同时设计形态树扩散方法(Morphological Tree Diffusion),强化不同末端执行器间的动作行为关联,最大化异构示范的可迁移性。在RoboCasa与Isaac Gym上的实验结果表明,X-DiffVLA在多种机器人形态(从夹爪到灵巧手)下均达当前最优,性能提升分别为15.3%和12.5%。真实世界测试进一步验证了该框架的鲁棒性及其在可扩展跨形态策略学习中的有效性。

原文摘要 · Abstract (English)

Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-trained on large and diverse datasets, they typically rely on embodiment-specific fine-tuning to achieve strong performance in downstream tasks. This requirement severely limits their generalization capability and restricts knowledge transfer across embodiments performing similar tasks. To overcome these limitations, we focus on cross-embodied settings with shared robotic bases and heterogeneous end-effectors, and propose X-DiffVLA, a diffusion-based VLA model featuring a unified cross-embodied action head. X-DiffVLA can leverage the generative strengths of diffusion models to capture both the diversity and latent correlations in cross-embodied datasets. Specifically, we introduce Embodiment Forcing, a classifier-free guidance technique to implicitly steer action generation toward embodiment-specific functional components, capturing fine-grained structural nuances without explicit supervision. In addition, a Morphological Tree Diffusion approach is designed to strengthen behavioral correlations across diverse end-effectors, maximizing the transferability of heterogeneous demonstrations. Experimental results across RoboCasa and Isaac Gym, covering different embodiments from grippers to dexterous hands, show that X-DiffVLA achieves state-of-the-art performance, with improvements of 15.3% and 12.5%, respectively. Real-world evaluations further validate the robustness of the proposed framework and its effectiveness in scalable cross-embodied policy learning.

机器人扩散模型跨形态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。