arXiv:2511.22134cs.CVcs.RO2025-11被引 12

解决视觉语言动作模型推理与执行失衡问题,提升通用机器人代理的综合表现。

DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action

  • 通过分层数据筛选与双教师自适应蒸馏,分离推理与动作学习
  • 在8个基准上平均得分65.4,仿真环境成功率达61.0%
  • 提出可拆解评估维度的VLA Score,适合研究通用智能体的开发者

为构建具备强推理能力的通用视觉-语言-动作(VLA)模型,现有方法通常先训练专用VLA获取可靠操作技能,再融合标注数据与多模态数据恢复更广的推理能力。然而我们发现,微调后推理型VLA的动作性能常显著下降,称为动作退化。为此,我们提出DualVLA,通过精心设计的后训练策略增强动作表现,同时保留推理能力。首先引入双层数据修剪方法,剔除冗余具身推理,避免其干扰动作学习;进一步设计双教师自适应蒸馏策略,在不同数据域分配差异化监督信号,强化动作生成。为填补通用VLA评估空白,我们还提出VLA Score,将能力分解为推理、意图、动作与对齐四个维度进行细粒度评估。实验表明,DualVLA在SimplerEnv中平均成功率61.0%,在八个竞争性多模态基准上平均得分为65.4,展现了精准动作执行与多模态理解之间的更强平衡。

原文摘要 · Abstract (English)

To build a generalizable Vision-Language-Action (VLA) model with strong reasoning ability, a common strategy is to first train a specialist VLA on robot demonstrations to acquire reliable manipulation skills, and then incorporate mixed annotated robot data together with multimodal data to restore broader reasoning capabilities. However, we observe that the resulting reasoning VLA often suffers from degraded action performance compared to the specialist model before fine-tuning, a phenomenon we refer to as action degeneration. To address this issue, we propose DualVLA, which enhances action performance through carefully designed post-training while still preserving reasoning capability. We first introduce a dual-layer data pruning method that removes redundant embodied reasoning, preventing it from adversely influencing action learning. To further strengthen action generation, we design a dual-teacher adaptive distillation strategy that assigns different supervision signals to different data domains while maintaining reasoning ability. To fill the evaluation gap for generalist VLAs, we also propose VLA Score, which decouples VLA capability into reasoning, intention, action, and alignment dimensions for a more fine-grained assessment. Experiments show that DualVLA achieves an average success rate of 61.0 in SimplerEnv and an average score of 65.4 across eight competitive multimodal benchmarks, demonstrating a stronger balance between precise action execution and multimodal understanding. Project Website: https://costaliya.github.io/DualVLA/.

通用智能体视觉语言动作机器人学习模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。