让视觉语言动作模型分离概念与操作知识,实现对新物体的零样本技能迁移。
Decoupling the Declarative from the Procedural in Vision-Language-Action Models

- 重构信息流,将语义与操作知识分离开来处理
- 在未见过的物体上实现前所未有的零样本技能迁移
- 适合需要强泛化能力的机器人通用智能研究
在真实世界部署通用机器人代理需要可迁移的技能。具体而言,从特定物体示范中学习行为的策略必须能推广到其他物体,否则数据收集成本会变得不可控。近期,基于大规模机器人数据集预训练、再用少量场景特定示范微调的千亿参数视觉-语言模型(VLM),已成为视觉-语言-动作(VLA)模型设计的主流范式。尽管这些策略在分布内任务中达到顶尖表现,但对空间、语义和任务的微小变化仍显脆弱。本文针对当前模型无法将陈述性知识(如概念与实体语义)与程序性知识(如何执行某事)解耦的问题,提出w²VLA新模型。该模型通过组合式、可解释的方式,以视觉、空间和技能信息调制机器人状态序列,而非将所有多模态编码令牌输入大型封闭式动作专家。相比现有先进VLA模型,我们的模块化方法成功解耦知识表征,实现稳健的行为克隆和跨不同未见物体的空前零样本技能迁移能力。
原文摘要 · Abstract (English)
Deploying generalist robotic agents in the real world requires transferable skills. Specifically, a policy trained to clone a behavior from object-specific demonstrations must generalize beyond that object, otherwise data collection requirements become intractable. Recently, fine-tuning of pre-trained billion-parameter Vision-Language Models (VLMs), initially on large-scale robot datasets and then on fewer scenario-specific demonstrations, has emerged as the predominant paradigm for designing Vision-Language-Action (VLA) models. While these policies achieve state-of-the-art manipulation performance in-distribution, they remain brittle to minor spatial, semantic, and task variations. In this work, we address the inability of current models to decouple the declarative (i.e., concepts and entity semantics) from the procedural knowledge (i.e., how to do something) encoded in their parameters, which is a fundamental bottleneck for zero-shot skill transfer to novel objects. To address this, we propose w$^{2}$VLA, a new VLA model with restructured information flow. Rather than feeding all multimodal tokens from the VLM encoder into a large, opaque transformer-based action expert, our approach modulates the robot state sequence with visual, spatial, and skill information in a compositional and interpretable manner. Unlike popular, state-of-the-art VLAs, we show that our modular approach successfully decouples knowledge representations, enabling robust behavior cloning and unprecedented zero-shot skill transfer capabilities across dissimilar, unseen objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。