提出高效视觉令牌压缩方法,提升高分辨率文本导向视觉语言模型的推理效率。
FCoT-VL:Advancing Text-oriented Large Vision-Language Models with Efficient Visual Token Compression
- 轻量级自蒸馏预训练压缩视觉令牌,仅需少量图文对和极少可学习参数。
- 在多文本导向基准测试中,压缩后模型性能优于基线,计算开销显著降低。
- 适合追求高分辨率图像理解与部署效率的研究者和开发者使用。
视觉大语言模型(VLLMs)的快速发展常依赖于包含丰富视觉令牌的高分辨率图像,这制约了训练与部署效率。现有无训练的视觉令牌压缩方法在涉及高分辨率、文本导向的图像理解与推理任务中表现严重退化。本文针对高分辨率场景下文本导向的VLLMs,提出一种高效的视觉令牌压缩框架。具体而言,采用轻量级自蒸馏预训练阶段压缩视觉令牌,仅需少量图像-文本对和极少可学习参数。随后,为缓解压缩模型潜在的性能下降,构建高质量后训练阶段。通过应用于先进模型InternVL2的实验表明,该方法显著降低计算开销,同时在多个文本导向基准上超越基线。相关模型与代码即将开源。
原文摘要 · Abstract (English)
The rapid success of Vision Large Language Models (VLLMs) often depends on the high-resolution images with abundant visual tokens, which hinders training and deployment efficiency. Current training-free visual token compression methods exhibit serious performance degradation in tasks involving high-resolution, text-oriented image understanding and reasoning. In this paper, we propose an efficient visual token compression framework for text-oriented VLLMs in high-resolution scenarios. In particular, we employ a light-weight self-distillation pre-training stage to compress the visual tokens, requiring a limited numbers of image-text pairs and minimal learnable parameters. Afterwards, to mitigate potential performance degradation of token-compressed models, we construct a high-quality post-train stage. To validate the effectiveness of our method, we apply it to an advanced VLLMs, InternVL2. Experimental results show that our approach significantly reduces computational overhead while outperforming the baselines across a range of text-oriented benchmarks. We will release the models and code soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。