用压缩视觉信息提升效率,让多模态模型跑得更快
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
- 将图像压缩与关键信息提取结合,减少冗余视觉标记
- 视觉标记数、训练内存和计算量显著降低,性能仍保持领先
- 适合需要快速响应的轻量级多模态应用
尽管多模态大语言模型(MLLMs)能力大幅提升,但在实际应用中仍像树懒般响应缓慢、延迟高。现有研究致力于构建小型化MLLM以提升效率,但大量视觉标记仍限制实际提速。本文提出一种高效且快速的小型化多模态大模型FlashSloth。不同于以往方法,FlashSloth在压缩视觉标记冗余语义的同时增强其描述能力,引入嵌入式视觉压缩设计,有效捕获图像中视觉显著和指令相关的信息,从而以更少的视觉标记实现优异的多模态表现。通过大量实验验证,相较于InternVL2、MiniCPM-V2和Qwen2-VL等先进小型化MLLMs,FlashSloth显著减少视觉标记数量、训练内存消耗和计算复杂度,同时在多种视觉语言任务上保持高精度表现。
原文摘要 · Abstract (English)
Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent efforts are devoted to building tiny MLLMs for better efficiency, but the plethora of visual tokens still used limit their actual speedup. In this paper, we propose a powerful and fast tiny MLLM called FlashSloth. Different from previous efforts, FlashSloth focuses on improving the descriptive power of visual tokens in the process of compressing their redundant semantics. In particular, FlashSloth introduces embedded visual compression designs to capture both visually salient and instruction-related image information, so as to achieving superior multimodal performance with fewer visual tokens. Extensive experiments are conducted to validate the proposed FlashSloth, and a bunch of tiny but strong MLLMs are also comprehensively compared, e.g., InternVL2, MiniCPM-V2 and Qwen2-VL. The experimental results show that compared with these advanced tiny MLLMs, our FlashSloth can greatly reduce the number of visual tokens, training memory and computation complexity while retaining high performance on various VL tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。