arXiv:2502.06788cs.CVcs.AI2025-02ICCV被引 34

提出更高效的无编码器视觉语言模型,性能逼近主流模型。

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

  • 通过层次化关联视觉与语言模态,减少干扰
  • 设计优化训练策略,实现高效学习
  • 适合追求轻量部署与多模态推理的开发者

现有无编码器视觉语言模型(VLMs)正快速缩小与有编码器模型的性能差距,展现出结构简化与高效部署的统一多模态系统潜力。本文系统厘清了使用预训练视觉编码器、离散分词器及从零开始的极简视觉层的VLMs之间的性能差异,深入挖掘无编码器VLMs的未被充分研究特性。我们提出高效策略,使无编码器VLMs性能媲美主流编码器模型。经深入探究,推出EVEv2.0——新一代改进型无编码器VLM家族。实验表明:(i) 在统一模型中合理分解并层次化关联视觉与语言,可降低模态间干扰;(ii) 精心设计的训练策略可有效优化无编码器VLM。通过广泛评估,EVEv2.0全面验证了跨模态解码器仅架构的可行性,展现出优异的数据效率与强视觉推理能力。代码公开于:https://github.com/baaivision/EVE。

原文摘要 · Abstract (English)

Existing encoder-free vision-language models (VLMs) are rapidly narrowing the performance gap with their encoder-based counterparts, highlighting the promising potential for unified multimodal systems with structural simplicity and efficient deployment. We systematically clarify the performance gap between VLMs using pre-trained vision encoders, discrete tokenizers, and minimalist visual layers from scratch, deeply excavating the under-examined characteristics of encoder-free VLMs. We develop efficient strategies for encoder-free VLMs that rival mainstream encoder-based ones. After an in-depth investigation, we launch EVEv2.0, a new and improved family of encoder-free VLMs. We show that: (i) Properly decomposing and hierarchically associating vision and language within a unified model reduces interference between modalities. (ii) A well-designed training strategy enables effective optimization for encoder-free VLMs. Through extensive evaluation, our EVEv2.0 represents a thorough study for developing a decoder-only architecture across modalities, demonstrating superior data efficiency and strong vision-reasoning capability. Code is publicly available at: https://github.com/baaivision/EVE.

视觉语言模型无编码器多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。