arXiv:2503.10665cs.CVcs.AI2025-03综述被引 5

小规模多模态模型让高效智能在低资源环境成为可能

Small Vision-Language Models: A Survey on Compact Architectures and Techniques

  • 按架构分类:基于Transformer、Mamba及混合结构的紧凑设计
  • 通过轻量化注意力与知识蒸馏,实现低资源下的高精度表现
  • 适合边缘计算与移动端部署,关注模型效率与泛化能力

小型视觉语言模型(sVLMs)的出现标志着多模态AI的重要进展,可在资源受限环境中高效处理视觉与文本数据。本文系统梳理sVLM的发展,提出架构分类体系:基于Transformer、Mamba及混合架构,凸显紧凑设计与计算效率的创新。讨论知识蒸馏、轻量级注意力机制与模态预融合等关键技术,推动在降低资源消耗的同时保持高性能。通过对TinyGPT-V、MiniGPT-4和VL-Mamba等模型的深入分析,揭示准确率、效率与可扩展性之间的权衡。针对数据偏差与复杂任务泛化能力不足等持续挑战,提出潜在解决路径。本综述整合了sVLM的关键进展,凸显其在推动普惠AI中的变革潜力,为未来高效多模态系统研究奠定基础。

原文摘要 · Abstract (English)

The emergence of small vision-language models (sVLMs) marks a critical advancement in multimodal AI, enabling efficient processing of visual and textual data in resource-constrained environments. This survey offers a comprehensive exploration of sVLM development, presenting a taxonomy of architectures - transformer-based, mamba-based, and hybrid - that highlight innovations in compact design and computational efficiency. Techniques such as knowledge distillation, lightweight attention mechanisms, and modality pre-fusion are discussed as enablers of high performance with reduced resource requirements. Through an in-depth analysis of models like TinyGPT-V, MiniGPT-4, and VL-Mamba, we identify trade-offs between accuracy, efficiency, and scalability. Persistent challenges, including data biases and generalization to complex tasks, are critically examined, with proposed pathways for addressing them. By consolidating advancements in sVLMs, this work underscores their transformative potential for accessible AI, setting a foundation for future research into efficient multimodal systems.

多模态轻量化模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。