让视觉语言模型更快更稳,适合实时应用部署。
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
- 用架构优化与语义拼接技术降低计算负担。
- 11项基准测试中速度与准确率均达顶尖水平。
- 适合边缘设备和大规模实时系统使用。
本文提出Flash-VL 2B,一种面向实时应用的视觉语言模型优化方法,旨在实现超低延迟与高吞吐量,同时保持准确率。通过定制化架构设计、令牌压缩机制、数据筛选、训练策略以及一种名为隐式语义拼接的新图像处理技术,有效平衡计算负载与性能。在11个标准视觉语言模型基准上进行广泛评估,结果表明该模型在速度与准确率方面均达到领先水平,是资源受限环境与大规模实时应用场景的理想选择。
原文摘要 · Abstract (English)
In this paper, we introduce Flash-VL 2B, a novel approach to optimizing Vision-Language Models (VLMs) for real-time applications, targeting ultra-low latency and high throughput without sacrificing accuracy. Leveraging advanced architectural enhancements and efficient computational strategies, Flash-VL 2B is designed to maximize throughput by reducing processing time while maintaining competitive performance across multiple vision-language benchmarks. Our approach includes tailored architectural choices, token compression mechanisms, data curation, training schemes, and a novel image processing technique called implicit semantic stitching that effectively balances computational load and model performance. Through extensive evaluations on 11 standard VLM benchmarks, we demonstrate that Flash-VL 2B achieves state-of-the-art results in both speed and accuracy, making it a promising solution for deployment in resource-constrained environments and large-scale real-time applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。