arXiv:2511.21542cs.ROcs.AI2025-11被引 3

用扩散模型生成更精准的机器人动作,提升跨场景泛化能力

E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion

  • 通过离散扩散过程迭代去噪生成动作令牌
  • 在4个数据集上平均性能领先基线10.7%
  • 适合需要精细控制的机器人任务研究者

视觉-语言-动作(VLA)模型通过融合视觉感知、语言理解与控制生成,为机器人操作提供统一框架。然而现有系统在任务、场景和相机视角间泛化能力弱,常产生粗粒度或不稳定的动作。我们指出这些局限与动作分布的多峰特性、预训练模型的符号化推理机制及真实机器人控制的有限分辨率密切相关。为此,提出E0——一种基于Tweedie离散扩散的框架,将动作生成建模为量化动作令牌上的迭代去噪过程。该方法在离散动作空间中以合理扩散流程运行,天然适配符号推理,支持细粒度且可执行的动作控制,并避免掩码式离散扩散的分布错位问题。进一步引入球面视角扰动增强策略,在无需额外数据下提升对相机位移的鲁棒性。在LIBERO、VLABench、ManiSkill及真实Franka机械臂上的实验表明,E0在14种不同环境中均达到当前最优表现,平均超越强基线10.7%。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, existing VLA systems still struggle to generalize across diverse tasks, scenes, and camera viewpoints, and often produce coarse or unstable actions. We argue that these limitations are closely tied to the structural properties of actions in VLA settings, including the inherent multi-peaked nature of action distributions, the token-based symbolic reasoning of pretrained VLM/VLA backbones, and the effective finite resolution imposed by real-world robotic control. Motivated by these properties, we introduce E0, a tweedie discrete diffusion framework that formulates action generation as iterative denoising over quantized action tokens. By operating in a discrete action space with a principled diffusion process, E0 naturally aligns with token-based reasoning, supports fine-grained yet executable action control, and avoids the distributional mismatch of masking-based discrete diffusion. We further introduce a spherical viewpoint perturbation augmentation to enhance robustness to camera shifts without additional data. Experiments on LIBERO, VLABench, ManiSkill, and a real-world Franka arm demonstrate that E0 achieves state-of-the-art performance across 14 diverse environments, outperforming strong baselines by 10.7% on average.

机器人控制扩散模型动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。