通过表情符号任务测试视觉语言模型知识如何迁移到机器人控制。
How Do VLAs Effectively Inherit from VLMs?
- 设计表情符号操作任务,检验视觉语言模型先验知识能否迁移到机器人
- 实验证明保留视觉语言模型先验对泛化能力至关重要
- 适合研究通用具身智能与知识迁移的学者参考
视觉-语言-动作(VLA)模型有望实现可泛化的具身控制。主流方法是利用大规模视觉-语言模型(VLM)丰富的视觉语义先验。然而根本问题仍存:VLA如何有效继承VLM的知识?为此,我们提出诊断基准GrinningFace——一个表情符号桌面操作任务,要求机械臂根据语言指令将物体放置在印刷的表情符号上。该任务设计具有揭示性:表情符号在互联网规模数据集广泛存在,但几乎未见于标准机器人数据集,因此可作为清晰代理指标——任务成功完成即表明VLM先验有效转移至具身控制。我们在仿真环境和真实机器人上实现该任务,对比参数高效微调、冻结VLM、联合训练、离散动作预测及潜在动作预测等多种知识迁移技术。系统评估表明,保留VLM先验对VLA泛化至关重要,并为未来构建真正通用的具身智能系统提供指导。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models hold the promise to attain generalizable embodied control. To achieve this, a pervasive paradigm is to leverage the rich vision-semantic priors of large vision-language models (VLMs). However, the fundamental question persists: How do VLAs effectively inherit the prior knowledge from VLMs? To address this critical question, we introduce a diagnostic benchmark, GrinningFace, an emoji tabletop manipulation task where the robot arm is asked to place objects onto printed emojis corresponding to language instructions. This task design is particularly revealing -- knowledge associated with emojis is ubiquitous in Internet-scale datasets used for VLM pre-training, yet emojis themselves are largely absent from standard robotics datasets. Consequently, they provide a clean proxy: successful task completion indicates effective transfer of VLM priors to embodied control. We implement this diagnostic task in both simulated environment and a real robot, and compare various promising techniques for knowledge transfer. Specifically, we investigate the effects of parameter-efficient fine-tuning, VLM freezing, co-training, predicting discretized actions, and predicting latent actions. Through systematic evaluation, our work not only demonstrates the critical importance of preserving VLM priors for the generalization of VLA but also establishes guidelines for future research in developing truly generalizable embodied AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。