arXiv:2503.14694cs.CLcs.CV2025-03ICML被引 10

用单个Transformer实现多模态理解,性能逼近主流模型。

HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding

  • 早期融合视觉与文本输入,统一建模于单个Transformer中。
  • 训练方法有效利用预训练知识,性能超越同类单模型。
  • 适合追求高效端到端多模态系统的研究者参考。

近期大型语言模型(LLMs)的发展推动了大规模多模态模型(LMMs)的进步,展现出通用智能助手的潜力。然而,多数LMMs对视觉和文本模态分别建模,促使研究者探索基于单一Transformer的原生多模态模型。尽管前景广阔,这些原生模型通常资源消耗大,且性能落后于组合式模型。为此,我们提出一种简单高效的单变压器基线方法,构建原生端到端的大规模多模态模型。首先,设计一种新的早期融合LMM,能在早期阶段融合多模态输入,并以自回归方式响应视觉指令。其次,制定一种高效的训练策略,利用预训练模型的先验知识,缓解性能瓶颈与资源消耗问题。所提模型在单变压器架构下表现更优,显著缩小了与组合式模型之间的性能差距。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have significantly propelled the development of large multi-modal models (LMMs), highlighting the potential for general and intelligent assistants. However, most LMMs model visual and textual modalities separately, leading to recent efforts to develop native LMMs using a single transformer. Despite the promise, these native models are resource-intensive and often exhibit performance gaps compared to their compositional counterparts. To alleviate this issue, we propose a simple yet efficient method to construct a baseline for the native and end-to-end large multi-modal model in a single transformer. First, we propose a new early-fusion LMM that can fuse multi-modal inputs in the early stage and respond to visual instructions in an auto-regressive manner. Second, we devise an efficient training recipe for the proposed model, which harnesses the prior knowledge of the pre-trained models, addressing both the performance limitations and the challenge of resource consumption. The proposed model demonstrates superior performance compared to other LMMs using one transformer and significantly narrows the performance gap with compositional LMMs.

多模态单变压器视觉语言基线模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。