将流模型视觉语言动作模型深度从18层压缩到6层,实现更快推理且性能损失极小。
Shallow-π: Knowledge Distillation for Flow-based VLAs
- 通过知识蒸馏大幅减少视觉语言动作模型的变压器层数
- 推理速度提升两倍以上,成功率仅下降不到1%
- 在真实机器人平台上验证,适合工业级实时部署
实时机器人部署对视觉-语言-动作(VLA)模型的快速、本地化推理需求日益增长。现有研究多聚焦于令牌层面的效率优化,如视觉令牌剪枝,但对变压器层数的系统性缩减关注较少,尤其在基于流的VLA模型中尚未探索知识蒸馏方法。本文提出Shallow-π,一种有原则的知识蒸馏框架,显著压缩视觉语言模型主干和基于流的动作头的深度,将模型层数从18层降至6层。该方法在标准操作基准上实现超过两倍的推理加速,成功率绝对下降不足1%,在压缩后的VLA模型中达到当前最优表现。关键的是,我们在Jetson Orin与Jetson Thor平台上,跨多个机器人平台(包括人形机器人)的复杂动态操作场景中,进行了大规模真实世界实验,验证了方法的有效性。
原文摘要 · Abstract (English)
The growing demand for real-time robotic deployment necessitates fast and on-device inference for vision-language-action (VLA) models. Within the VLA literature, efficiency has been extensively studied at the token level, such as visual token pruning. In contrast, systematic transformer layer reduction has received limited attention and, to the best of our knowledge, has not been explored for flow-based VLA models under knowledge distillation. In this work, we propose Shallow-pi, a principled knowledge distillation framework that aggressively reduces the transformer depth of both the VLM backbone and the flow-based action head, compressing the model from 18 to 6 layers. Shallow-pi achieves over two times faster inference with less than one percent absolute drop in success rate on standard manipulation benchmarks, establishing state-of-the-art performance among reduced VLA models. Crucially, we validate our approach through industrial-scale real-world experiments on Jetson Orin and Jetson Thor across multiple robot platforms, including humanoid systems, in complex and dynamic manipulation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。