用指令引导压缩视觉标记,让机器人更高效精准地执行任务
Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
- 根据自然语言指令动态压缩视觉信息,保留关键上下文和细节
- 在LIBERO基准上成功率达90%以上,计算量降低59%,视觉标记减少3倍以上
- 适合需要实时推理的机器人控制场景,尤其对多臂操作有实用价值
视觉-语言-动作(VLA)模型在具身智能中表现出强大能力,但冗余视觉标记带来的高计算开销仍是实时机器人部署的主要瓶颈。现有通用标记剪枝方法难以保留任务相关的关键视觉信息。为此,我们提出Compressor-VLA,一种新型混合式指令引导视觉标记压缩框架,实现任务导向的高效压缩。该框架包含两个模块:语义任务压缩器(STC)提取全局任务相关上下文,空间精修压缩器(SRC)保留精细空间细节。压缩过程由自然语言指令动态调控,实现任务相关的感知聚焦。实验表明,Compressor-VLA在LIBERO基准上保持90%以上的成功率,同时将浮点运算量(FLOPs)降低59%,视觉标记数量减少3倍以上。双臂机器人平台上的真实部署验证了其从仿真到现实的迁移能力与实际可用性。定性分析显示,指令引导有效引导模型关注任务相关物体,证实了方法的有效性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as a powerful paradigm in Embodied AI. However, the significant computational overhead of processing redundant visual tokens remains a critical bottleneck for real-time robotic deployment. While standard token pruning techniques can alleviate this, these task-agnostic methods struggle to preserve task-critical visual information. To address this challenge, simultaneously preserving both the holistic context and fine-grained details for precise action, we propose Compressor-VLA, a novel hybrid instruction-conditioned token compression framework designed for efficient, task-oriented compression of visual information in VLA models. The proposed Compressor-VLA framework consists of two token compression modules: a Semantic Task Compressor (STC) that distills holistic, task-relevant context, and a Spatial Refinement Compressor (SRC) that preserves fine-grained spatial details. This compression is dynamically modulated by the natural language instruction, allowing for the adaptive condensation of task-relevant visual information. Experimentally, extensive evaluations demonstrate that Compressor-VLA achieves a competitive success rate on the LIBERO benchmark while reducing FLOPs by 59% and the visual token count by over 3x compared to its baseline. The real-robot deployments on a dual-arm robot platform validate the model's sim-to-real transferability and practical applicability. Moreover, qualitative analyses reveal that our instruction guidance dynamically steers the model's perceptual focus toward task-relevant objects, thereby validating the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。