Pheye架构让高分辨率图文模型更高效,细粒度理解更强。
Efficient Architectures for High Resolution Vision-Language Models
- 采用新架构,处理高分辨率图像更高效
- 参数量少于同类模型,仍保持强性能
- 适合需要精细图像理解的任务
视觉语言模型(VLMs)近年来取得显著进展,但在高分辨率图像中准确识别细微细节方面仍面临挑战,限制了其在多个任务中的表现。本文提出Pheye,一种新型架构,在处理高分辨率图像时具有高效率,同时训练参数量少于同等规模的VLMs。Pheye在需要细粒度图像理解及场景文本处理的任务中表现出色,兼具高效性与高性能。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have recently experienced significant advancements. However, challenges persist in the accurate recognition of fine details within high resolution images, which limits performance in multiple tasks. This work introduces Pheye, a novel architecture that efficiently processes high-resolution images while training fewer parameters than similarly sized VLMs. Notably, Pheye achieves a high efficiency while maintaining strong performance, particularly in tasks that demand fine-grained image understanding and/or the handling of scene-text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。