arXiv:2502.07855cs.CVcs.AI2025-02综述被引 72

如何让大模型在手机、摄像头等边缘设备上高效运行

Vision-Language Models for Edge Networks: A Comprehensive Survey

  • 用剪枝、量化等技术压缩大模型,降低资源消耗
  • 提出轻量级训练与部署方案,支持实时视频分析
  • 适合做智能安防、医疗监测的边缘AI开发者

视觉语言模型(VLMs)融合视觉理解与自然语言处理,可完成图像描述、视觉问答和视频分析等任务。尽管在自动驾驶、智能监控和医疗等领域表现优异,但其在计算能力、内存和能耗受限的边缘设备上部署仍面临挑战。本文综述了近期针对边缘环境优化VLMs的进展,重点包括剪枝、量化、知识蒸馏等模型压缩技术,以及专用硬件解决方案以提升效率。系统讨论了高效训练与微调方法、边缘部署难题及隐私问题。此外,还介绍了轻量级VLM在医疗、环境监测和自主系统中的多样化应用,展示其日益增长的影响。通过提炼关键设计策略、现有挑战并提出未来方向建议,旨在推动VLM在资源受限场景下的实用化部署。

原文摘要 · Abstract (English)

Vision Large Language Models (VLMs) combine visual understanding with natural language processing, enabling tasks like image captioning, visual question answering, and video analysis. While VLMs show impressive capabilities across domains such as autonomous vehicles, smart surveillance, and healthcare, their deployment on resource-constrained edge devices remains challenging due to processing power, memory, and energy limitations. This survey explores recent advancements in optimizing VLMs for edge environments, focusing on model compression techniques, including pruning, quantization, knowledge distillation, and specialized hardware solutions that enhance efficiency. We provide a detailed discussion of efficient training and fine-tuning methods, edge deployment challenges, and privacy considerations. Additionally, we discuss the diverse applications of lightweight VLMs across healthcare, environmental monitoring, and autonomous systems, illustrating their growing impact. By highlighting key design strategies, current challenges, and offering recommendations for future directions, this survey aims to inspire further research into the practical deployment of VLMs, ultimately making advanced AI accessible in resource-limited settings.

视觉语言模型边缘计算模型压缩AI落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。