arXiv:2505.04601cs.CV2025-05ICCV被引 16

开源低成本视觉编码器,性能媲美CLIP,适配多模态模型部署

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning

论文配图:OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
图 1 · 摘自论文原文
  • 基于公开数据与训练框架构建,全流程开源
  • 590万至6.32亿参数模型,支持高效多模态应用
  • 适合需轻量部署或高性能的多模态研究者使用

OpenAI的CLIP自2021年初发布以来,长期是构建多模态基础模型的首选视觉编码器。尽管近期如SigLIP等替代方案开始挑战这一地位,但据我们所知,目前尚无完全开源的方案:其训练数据仍为专有,或训练方法未公开。本文提出OpenVision,一个完全开源、成本低廉的先进视觉编码器系列,在集成至LLaVA等多模态框架时,性能可达到甚至超越OpenAI的CLIP。OpenVision基于现有工作(如CLIPS训练框架和Recap-DataComp-1B训练数据),揭示了提升编码器质量的关键洞见,并展示了在推进多模态模型方面的实际效益。通过发布参数量从590万到6.32亿不等的编码器,OpenVision为实践者提供了在模型容量与效率之间灵活权衡的能力:大模型带来更强的多模态表现,小模型则支持轻量级、边缘就绪的多模态部署。

原文摘要 · Abstract (English)

OpenAI's CLIP, released in early 2021, have long been the go-to choice of vision encoder for building multimodal foundation models. Although recent alternatives such as SigLIP have begun to challenge this status quo, to our knowledge none are fully open: their training data remains proprietary and/or their training recipes are not released. This paper fills this gap with OpenVision, a fully-open, cost-effective family of vision encoders that match or surpass the performance of OpenAI's CLIP when integrated into multimodal frameworks like LLaVA. OpenVision builds on existing works -- e.g., CLIPS for training framework and Recap-DataComp-1B for training data -- while revealing multiple key insights in enhancing encoder quality and showcasing practical benefits in advancing multimodal models. By releasing vision encoders spanning from 5.9M to 632.1M parameters, OpenVision offers practitioners a flexible trade-off between capacity and efficiency in building multimodal models: larger models deliver enhanced multimodal performance, while smaller versions enable lightweight, edge-ready multimodal deployments.

视觉编码器多模态开源模型轻量化部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。