arXiv:2505.21567cs.CVcs.LG2025-05被引 2

针对视觉语言动作模型的对齐难题,提出编码对齐量化方法,显著降低计算开销。

EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models

  • 设计编码对齐分析法,定位多粒度下的特征对齐偏差
  • 提出感知编码对齐的混合精度量化,实现最小量化损失与加速
  • 适合追求高效推理的机器人控制与端到端智能体应用

随着具身人工智能的发展,端到端控制策略如视觉语言动作(VLA)模型已成为主流。现有VLA模型面临高昂的计算与存储成本,亟需优化。量化被视为最有效的方法,既能降低内存占用,又能加速计算。然而,我们发现VLA模型中的标记对齐问题阻碍了现有量化方法的应用。为此,提出优化框架EaqVLA,采用编码对齐量化。具体地,提出完整的分析方法以识别不同粒度下的错位情况;基于分析结果,设计考虑编码对齐意识的混合精度量化。实验表明,所提EaqVLA在端到端动作控制中实现最小量化损失,并获得xxx倍加速,优于现有量化方法。

原文摘要 · Abstract (English)

With the development of Embodied Artificial intelligence, the end-to-end control policy such as Vision-Language-Action (VLA) model has become the mainstream. Existing VLA models faces expensive computing/storage cost, which need to be optimized. Quantization is considered as the most effective method which can not only reduce the memory cost but also achieve computation acceleration. However, we find the token alignment of VLA models hinders the application of existing quantization methods. To address this, we proposed an optimized framework called EaqVLA, which apply encoding-aligned quantization to VLA models. Specifically, we propose an complete analysis method to find the misalignment in various granularity. Based on the analysis results, we propose a mixed precision quantization with the awareness of encoding alignment. Experiments shows that the porposed EaqVLA achieves better quantization performance (with the minimal quantization loss for end-to-end action control and xxx times acceleration) than existing quantization methods.

量化视觉语言动作模型压缩具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。