arXiv:2602.07849cs.AI2026-02

轻量级框架让视觉语言模型在边缘设备上高效自适应。

LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge

  • 用模态感知量化+无梯度自适应,兼顾性能与资源消耗。
  • 在七大数据集上内存降低19.9倍,适应性提升4.5%。
  • 适合隐私敏感、算力受限的边缘部署场景。

将视觉语言模型(VLMs)部署于边缘设备面临资源限制和分布偏移导致的性能下降问题。尽管测试时自适应(TTA)能缓解偏移影响,但现有方法对设备资源要求过高。为此,我们提出LQA——一种轻量级、量化自适应框架,融合模态感知量化策略与无梯度测试时自适应机制。引入选择性混合量化(SHQ)与量化无梯度适应方法,实现资源受限硬件上的鲁棒高效部署。在合成与真实世界分布偏移下实验表明,LQA整体适应性能提升4.5%,内存占用低于全精度模型,并显著优于基于梯度的TTA方法,在七个开源数据集上内存使用最高降低19.9倍。结果证明LQA为边缘设备上鲁棒、隐私保护、高效的VLM部署提供了可行路径。

原文摘要 · Abstract (English)

Deploying Vision-Language Models (VLMs) on edge devices is challenged by resource constraints and performance degradation under distribution shifts. While test-time adaptation (TTA) can counteract such shifts, existing methods are too resource-intensive for on-device deployment. To address this challenge, we propose LQA, a lightweight, quantized-adaptive framework for VLMs that combines a modality-aware quantization strategy with gradient-free test-time adaptation. We introduce Selective Hybrid Quantization (SHQ) and a quantized, gradient-free adaptation mechanism to enable robust and efficient VLM deployment on resource-constrained hardware. Experiments across both synthetic and real-world distribution shifts show that LQA improves overall adaptation performance by 4.5\%, uses less memory than full-precision models, and significantly outperforms gradient-based TTA methods, achieving up to 19.9$\times$ lower memory usage across seven open-source datasets. These results demonstrate that LQA offers a practical pathway for robust, privacy-preserving, and efficient VLM deployment on edge devices.

边缘计算视觉语言模型量化自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。