arXiv:2504.09724cs.CV2025-04综述被引 47

综述高效视觉语言模型的压缩与优化技术,助力边缘设备部署。

A Survey on Efficient Vision-Language Models

  • 系统梳理VLM在资源受限设备上的压缩与加速方法
  • 分析不同架构在性能与内存间的权衡关系
  • 适合关注AI轻量化与边缘计算的研究者

视觉语言模型(VLMs)融合视觉与文本信息,广泛应用于图像描述生成和视觉问答等任务,是现代人工智能系统的关键组成部分。然而,其高计算需求限制了在实时应用中的部署。为此,研究者们日益关注高效VLM的开发。本文综述了面向边缘和资源受限设备优化VLM的关键技术,探讨紧凑型VLM架构与框架,并深入分析高效VLM在性能与内存之间的权衡。此外,我们在https://github.com/MPSCUMBC/Efficient-Vision-Language-Models-A-Survey建立了一个GitHub仓库,持续收集并更新所调研的论文,旨在推动该领域的深入研究。

原文摘要 · Abstract (English)

Vision-language models (VLMs) integrate visual and textual information, enabling a wide range of applications such as image captioning and visual question answering, making them crucial for modern AI systems. However, their high computational demands pose challenges for real-time applications. This has led to a growing focus on developing efficient vision language models. In this survey, we review key techniques for optimizing VLMs on edge and resource-constrained devices. We also explore compact VLM architectures, frameworks and provide detailed insights into the performance-memory trade-offs of efficient VLMs. Furthermore, we establish a GitHub repository at https://github.com/MPSCUMBC/Efficient-Vision-Language-Models-A-Survey to compile all surveyed papers, which we will actively update. Our objective is to foster deeper research in this area.

视觉语言模型模型压缩边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。