arXiv:2412.05271cs.CV2024-12被引 1.7k

开源多模态模型InternVL 2.5通过数据与测试扩展,性能媲美GPT-4o

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

  • 采用模型、数据与测试时扩展策略提升性能
  • 在MMMU上达70.3%准确率,比前代高3.7点
  • 适合研究多模态推理与开放生态应用的开发者

我们提出InternVL 2.5,一个基于InternVL 2.0架构的先进多模态大语言模型系列,在保持核心结构的基础上,显著优化了训练与测试策略及数据质量。本文系统探究了视觉编码器、语言模型、数据集规模与测试时配置对性能的影响。在涵盖跨学科推理、文档理解、多图像/视频理解、真实世界认知、多模态幻觉检测、视觉定位、多语言能力及纯语言处理等广泛基准上的评估显示,InternVL 2.5表现优异,可与GPT-4o和Claude-3.5-Sonnet等领先商用模型媲美。尤为关键的是,该模型首次实现开源多模态模型在MMMU基准上超过70%(达70.3%),通过链式思维(CoT)推理提升3.7个百分点,并展现出强大的测试时扩展潜力。我们希望此模型为开源社区树立多模态AI系统研发与应用的新标准。HuggingFace演示见https://huggingface.co/spaces/OpenGVLab/InternVL

原文摘要 · Abstract (English)

We introduce InternVL 2.5, an advanced multimodal large language model (MLLM) series that builds upon InternVL 2.0, maintaining its core model architecture while introducing significant enhancements in training and testing strategies as well as data quality. In this work, we delve into the relationship between model scaling and performance, systematically exploring the performance trends in vision encoders, language models, dataset sizes, and test-time configurations. Through extensive evaluations on a wide range of benchmarks, including multi-discipline reasoning, document understanding, multi-image / video understanding, real-world comprehension, multimodal hallucination detection, visual grounding, multilingual capabilities, and pure language processing, InternVL 2.5 exhibits competitive performance, rivaling leading commercial models such as GPT-4o and Claude-3.5-Sonnet. Notably, our model is the first open-source MLLMs to surpass 70% on the MMMU benchmark, achieving a 3.7-point improvement through Chain-of-Thought (CoT) reasoning and showcasing strong potential for test-time scaling. We hope this model contributes to the open-source community by setting new standards for developing and applying multimodal AI systems. HuggingFace demo see https://huggingface.co/spaces/OpenGVLab/InternVL

多模态模型开源模型测试时扩展视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。