arXiv:2501.04322cs.CV2025-01AAAI被引 18

Eve通过弹性视觉专家实现小模型高效多模态,兼顾语言与视觉能力。

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

  • 引入弹性视觉专家,在多阶段训练中动态适配视觉能力
  • 1.8B参数模型在语言和多模态任务上均显著优于同类小模型
  • 小于30亿参数时超越多数模型,适合边缘设备部署

多模态视觉语言模型(VLMs)随着模型规模和数据量的持续增长取得了显著进展。然而,在边缘设备上运行这些模型仍面临挑战。现有高效VLM方法常以牺牲语言能力为代价提升多模态性能,或需要大量训练。为此,我们提出高效多模态视觉语言模型框架Eve,通过在训练多个阶段灵活引入可适配的视觉专家,平衡了语言能力与多模态能力的保留与增强。该方法使仅18亿参数的Eve在多模态和语言任务上均表现优异。在参数量低于30亿的配置下,其在语言基准测试中表现突出,且在VLM基准测试中达到68.87%的领先水平;其多模态准确率甚至超过更大的70亿参数的LLaVA-1.5模型。代码已开源于https://github.com/rangmiao/Eve。

原文摘要 · Abstract (English)

Multimodal vision language models (VLMs) have made significant progress with the support of continuously increasing model sizes and data volumes. Running VLMs on edge devices has become a challenge for their widespread application. There are several efficient VLM efforts, but they often sacrifice linguistic capabilities to enhance multimodal abilities, or require extensive training. To address this quandary,we introduce the innovative framework of Efficient Vision Language Models with Elastic Visual Experts (Eve). By strategically incorporating adaptable visual expertise at multiple stages of training, Eve strikes a balance between preserving linguistic abilities and augmenting multimodal capabilities. This balanced approach results in a versatile model with only 1.8B parameters that delivers significant improvements in both multimodal and linguistic tasks. Notably, in configurations below 3B parameters, Eve distinctly outperforms in language benchmarks and achieves state-of-the-art results 68.87% in VLM Benchmarks. Additionally, its multimodal accuracy outstrips that of the larger 7B LLaVA-1.5 model. Our code is available at https://github.com/rangmiao/Eve.

多模态小模型边缘计算视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。