arXiv:2504.10479cs.CV2025-04被引 1.8k

InternVL3统一训练多模态与语言能力,性能超越开源模型。

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

  • 单阶段联合训练,同时学习文本与图像信息。
  • 780亿参数版在MMMU上达72.2分,领先开源模型。
  • 适合研究多模态大模型的开发者与学术团队。

我们提出InternVL3,是InternVL系列的重大升级,采用原生多模态预训练范式。不同于将纯文本大语言模型(LLM)改造为支持视觉输入的多模态大语言模型(MLLM),InternVL3在单一预训练阶段,通过多样化多模态数据和纯文本语料,联合获取多模态与语言能力。这一统一训练范式有效解决了传统后处理训练流程中的复杂性和对齐难题。为提升性能与可扩展性,InternVL3引入可变视觉位置编码(V2PE)以支持更长的多模态上下文,采用监督微调(SFT)与混合偏好优化(MPO)等先进后训练技术,并结合测试时缩放策略及优化的训练基础设施。大量实证评估表明,InternVL3在多种多模态任务中表现优异。特别是,InternVL3-78B在MMMU基准上取得72.2分,创下开源多模态模型新纪录。其性能仍能与ChatGPT-4o、Claude 3.5 Sonnet、Gemini 2.5 Pro等领先专有模型比肩,同时保持强大的纯语言能力。为践行开放科学理念,我们将公开发布训练数据与模型权重,推动下一代多模态大模型的研究与发展。

原文摘要 · Abstract (English)

We introduce InternVL3, a significant advancement in the InternVL series featuring a native multimodal pre-training paradigm. Rather than adapting a text-only large language model (LLM) into a multimodal large language model (MLLM) that supports visual inputs, InternVL3 jointly acquires multimodal and linguistic capabilities from both diverse multimodal data and pure-text corpora during a single pre-training stage. This unified training paradigm effectively addresses the complexities and alignment challenges commonly encountered in conventional post-hoc training pipelines for MLLMs. To further improve performance and scalability, InternVL3 incorporates variable visual position encoding (V2PE) to support extended multimodal contexts, employs advanced post-training techniques such as supervised fine-tuning (SFT) and mixed preference optimization (MPO), and adopts test-time scaling strategies alongside an optimized training infrastructure. Extensive empirical evaluations demonstrate that InternVL3 delivers superior performance across a wide range of multi-modal tasks. In particular, InternVL3-78B achieves a score of 72.2 on the MMMU benchmark, setting a new state-of-the-art among open-source MLLMs. Its capabilities remain highly competitive with leading proprietary models, including ChatGPT-4o, Claude 3.5 Sonnet, and Gemini 2.5 Pro, while also maintaining strong pure-language proficiency. In pursuit of open-science principles, we will publicly release both the training data and model weights to foster further research and development in next-generation MLLMs.

多模态模型统一训练开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。