arXiv:2509.09090cs.CVcs.AI2025-09被引 15

首次实现视觉-语言-动作模型的量化与剪枝协同加速,性能不降反升。

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models

  • 设计量化感知的剪枝策略,让二者在压缩中协同工作。
  • 在不训练情况下实现1.93倍推理加速,平均成功率提升4.5%。
  • 适合部署高阶具身智能模型的轻量化场景。

视觉-语言-动作(VLA)模型在具身智能方面展现出前所未有的能力,但其巨大的计算与内存开销限制了实际部署。现有压缩方法通常单独进行量化或标记剪枝,因存在兼容性问题,难以同时实现两者以获得整体效率提升。本文提出SQAP-VLA,首个结构化、无需训练的VLA推理加速框架,首次实现最先进的量化与标记剪枝同步应用。通过协同设计量化与剪枝流程,提出可在激进量化模型上运行的量化感知剪枝准则,并优化量化器设计以增强剪枝效果。应用于标准VLA模型时,该方法显著提升计算效率与推理速度,同时保持核心性能,相较原模型实现×1.93的加速,平均成功率最高提升4.5%。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models exhibit unprecedented capabilities for embodied intelligence. However, their extensive computational and memory costs hinder their practical deployment. Existing VLA compression and acceleration approaches conduct quantization or token pruning in an ad-hoc manner but fail to enable both for a holistic efficiency improvement due to an observed incompatibility. This work introduces SQAP-VLA, the first structured, training-free VLA inference acceleration framework that simultaneously enables state-of-the-art quantization and token pruning. We overcome the incompatibility by co-designing the quantization and token pruning pipeline, where we propose new quantization-aware token pruning criteria that work on an aggressively quantized model while improving the quantizer design to enhance pruning effectiveness. When applied to standard VLA models, SQAP-VLA yields significant gains in computational efficiency and inference speed while successfully preserving core model performance, achieving a $\times$1.93 speedup and up to a 4.5\% average success rate enhancement compared to the original model.

VLA模型量化剪枝高效推理具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。