arXiv:2509.14033cs.CV2025-09被引 13

SAIL-VL2是开源多模态模型,性能领先且高效可用。

SAIL-VL2 Technical Report

  • 采用大规模数据筛选与渐进式训练框架提升模型能力
  • 在106个数据集上表现优异,多项推理任务达顶尖水平
  • 适合研究者和开发者用于构建高效开源多模态系统

我们推出SAIL-VL2,一个面向全面多模态理解与推理的开源视觉语言基础模型(LVM)。作为SAIL-VL的升级版,SAIL-VL2在2B和8B参数规模下,在多个图像与视频基准测试中达到当前最优性能,展现出从细粒度感知到复杂推理的强大能力。其有效性源于三大核心创新:首先,构建大规模数据清洗流水线,结合评分与过滤策略,优化了图文、OCR、问答及视频数据的质量与分布,提升训练效率;其次,采用渐进式训练框架,从预训练视觉编码器(SAIL-ViT)出发,经多模态预训练,最终通过思维融合的SFT-RL混合范式系统增强模型能力;第三,架构创新突破密集LLM限制,引入高效的稀疏专家混合(MoE)设计。凭借这些贡献,SAIL-VL2在106个数据集上表现优异,在挑战性推理基准如MMMU和MathVista上取得领先结果。此外,在OpenCompass排行榜上,SAIL-VL2-2B在4B参数规模的开源模型中排名第一,为开源多模态社区提供高效可扩展的基础模型。

原文摘要 · Abstract (English)

We introduce SAIL-VL2, an open-suite vision-language foundation model (LVM) for comprehensive multimodal understanding and reasoning. As the successor to SAIL-VL, SAIL-VL2 achieves state-of-the-art performance at the 2B and 8B parameter scales across diverse image and video benchmarks, demonstrating strong capabilities from fine-grained perception to complex reasoning. Its effectiveness is driven by three core innovations. First, a large-scale data curation pipeline with scoring and filtering strategies enhances both quality and distribution across captioning, OCR, QA, and video data, improving training efficiency. Second, a progressive training framework begins with a powerful pre-trained vision encoder (SAIL-ViT), advances through multimodal pre-training, and culminates in a thinking-fusion SFT-RL hybrid paradigm that systematically strengthens model capabilities. Third, architectural advances extend beyond dense LLMs to efficient sparse Mixture-of-Experts (MoE) designs. With these contributions, SAIL-VL2 demonstrates competitive performance across 106 datasets and achieves state-of-the-art results on challenging reasoning benchmarks such as MMMU and MathVista. Furthermore, on the OpenCompass leaderboard, SAIL-VL2-2B ranks first among officially released open-source models under the 4B parameter scale, while serving as an efficient and extensible foundation for the open-source multimodal community.

多模态视觉语言开源模型推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。