arXiv:2607.02604cs.CVcs.RO2026-07

基于VLA模型的动态物体操作新框架,提升机器人抓取成功率。

DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation

论文配图:DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation
图 1 · 摘自论文原文
  • 用Mamba-3编码动作,融合视觉与本体感知信息生成动作轨迹。
  • 在多个基准上实现最高45.31%的成功率提升,尤其对微调模型效果显著。
  • 适合研究动态物体操控、强化学习与多模态机器人系统的人参考。

尽管视觉-语言-动作(VLA)模型受到广泛关注,但在操控动态移动物体方面仍面临诸多挑战。现有方法通常将世界模型嵌入高性能基础VLA架构中,可能导致预训练好的基础模型性能下降。本文提出DynaWM,一种基于基础VLA引导的世界基础模型,可适配多种微调和粗调的基础VLA检查点,用于动态物体操作。DynaWM采用Mamba-3-based动作编码器处理基础VLA输出的动作片段,生成动作条件表示;使用V-JEPA 2.1视觉编码器提取多视角观测历史特征,并通过本体感知编码器处理机械臂状态。这些特征联合控制一个流匹配的DiT模型,以重构具有运动感知的动作轨迹。为系统评估,构建了DynaGrasp-32基准,涵盖六类任务,包括速度变化、轨迹变化及多物体操作;以及包含32个场景、1,600条示范轨迹和约153万张图像的DynaGrasp-1600数据集。对于微调的基础VLA模型,DynaWM相较SmolVLA、X-VLA、π0、π0.5分别提升7.19%、45.31%、1.88%、10.94%;对粗调模型则分别提升35.13%、44.06%、35.69%、26.13%。消融实验表明,视觉编码使成功率提高27.50%,若移除动作条件则下降45.44%。

原文摘要 · Abstract (English)

Although vision-language-action (VLA) models have received widespread attention, many challenges remain in manipulating dynamic moving objects. In most existing approaches, end-to-end forward or inverse dynamics models, i.e., world models, are incorporated into high-performance base VLA architectures, which may degrade the performance of well-pretrained base VLA models due to inappropriate fine-tuning. In this paper, we propose DynaWM, a base-VLA-guided world foundation model that adapts to a wide variety of fine-tuned and coarse-tuned base-VLA checkpoints for moving-object manipulation. DynaWM uses a Mamba-3-based action encoder to encode the base action chunk produced by the base VLA into an action-conditioning representation, a V-JEPA 2.1 vision encoder to extract features from multi-view observation history, and a proprioceptive state encoder to encode robotic-arm proprioceptive states. These feature representations jointly condition a flow-matching DiT to regenerate motion-aware action trajectories for moving-object manipulation. For systematic evaluation, we construct the DynaGrasp-32 benchmark, covering six categories of moving-object manipulation tasks, including velocity variation, trajectory variation, and multi-object manipulation, as well as the DynaGrasp-1600 dataset, which consists of 32 scenarios, 1,600 demonstration trajectories, and approximately 1.53M images. For fine-tuned base-VLA checkpoints, DynaWM achieves percentage improvements of 7.19, 45.31, 1.88, and 10.94 over SmolVLA, X-VLA, π0, and π0.5, respectively. For coarse-tuned base-VLA checkpoints, performance increases by 35.13, 44.06, 35.69, and 26.13 percentage, respectively. Ablation experiments show that visual encoding enhances success by 27.50%, while reducing success by 45.44% if action conditioning is removed.

动态操作视觉语言机器人世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。