arXiv:2604.24182cs.RO2026-04

让视觉语言模型直接指挥机器人,通过分层混合与元技能提升泛化能力

$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills

论文配图:$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills
图 1 · 摘自论文原文
  • 用分层混合策略从语义特征中提取任务关键信息
  • 在有限模型容量下实现高效轨迹学习,零样本泛化效果好
  • 适合研究机器人操控与多模态模型融合的开发者

当前视觉-语言-动作(VLA)模型主要依赖端到端微调,虽有效但削弱了视觉-语言模型(VLM)的固有泛化能力,并引发灾难性遗忘。为此,我们提出 $M^2$-VLA,证明通用的 VLM 可直接作为机器人操作的强大主干网络。然而,如何弥合 VLM 的高层语义理解与机器人控制的精确需求之间的差距仍是关键挑战。为此,我们引入分层混合(MoL)策略,有选择地提取密集语义特征中的任务关键信息;同时,为在模型容量受限条件下实现高效轨迹学习,提出元技能模块(MSM),融入强归纳偏置。在模拟和真实世界环境中进行的大量实验验证了该方法的有效性。泛化与消融实验进一步证实了架构的零样本能力,并确认了各核心组件的贡献。代码与预训练模型将公开发布。

原文摘要 · Abstract (English)

Current Vision-Language-Action (VLA) models predominantly rely on end-to-end fine-tuning. While effective, this paradigm compromises the inherent generalization capabilities of Vision-Language Models (VLMs) and incurs catastrophic forgetting. To address these limitations, we propose $M^2$-VLA, which demonstrates that a generalized VLM is able to serve as a powerful backbone for robotic manipulation directly. However, it remains a key challenge to bridge the gap between the high-level semantic understanding of VLMs and the precise requirements of robotic control. To overcome this, we introduce the Mixture of Layers (MoL) strategy that selectively extracts task-critical information from dense semantic features. Furthermore, to facilitate efficient trajectory learning under constrained model capacity, we propose a Meta Skill Module (MSM) that integrates strong inductive biases. Extensive experiments in both simulated and real-world environments demonstrate the effectiveness of our approach. Furthermore, generalization and ablation studies validate the architecture's zero-shot capabilities and confirm the contribution of each key component. Our code and pre-trained models will be made publicly available.

机器人操控视觉语言模型元技能分层混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。