轻量级多模态模型在60个基准上38项领先,推理能力突出。
Seed1.5-VL Technical Report
- 采用532M视觉编码器与20B活跃参数的MoE语言模型
- 在60个公开基准中38项达到顶尖水平,GUI控制表现超越OpenAI CUA
- 擅长视觉推理任务,适合需要跨模态理解的应用场景
我们提出Seed1.5-VL,一个面向通用多模态理解与推理的视觉语言基础模型。该模型由532M参数的视觉编码器和20B活跃参数的混合专家(MoE)语言模型构成。尽管架构紧凑,其在广泛公共VLM基准及内部评估中均表现强劲,在60个公开基准中实现38项领先。在以代理为中心的任务(如GUI控制和游戏)中,其性能超越OpenAI CUA与Claude 3.7。此外,模型在视觉与视频理解之外展现出强大推理能力,尤其适用于视觉谜题等多模态推理挑战。本文重点回顾了模型设计、数据构建与训练各阶段的经验,旨在推动后续研究。Seed1.5-VL现已开放获取,可通过火山引擎平台(Volcano Engine Model ID: doubao-1-5-thinking-vision-pro-250428)访问。
原文摘要 · Abstract (English)
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter vision encoder and a Mixture-of-Experts (MoE) LLM of 20B active parameters. Despite its relatively compact architecture, it delivers strong performance across a wide spectrum of public VLM benchmarks and internal evaluation suites, achieving the state-of-the-art performance on 38 out of 60 public benchmarks. Moreover, in agent-centric tasks such as GUI control and gameplay, Seed1.5-VL outperforms leading multimodal systems, including OpenAI CUA and Claude 3.7. Beyond visual and video understanding, it also demonstrates strong reasoning abilities, making it particularly effective for multimodal reasoning challenges such as visual puzzles. We believe these capabilities will empower broader applications across diverse tasks. In this report, we mainly provide a comprehensive review of our experiences in building Seed1.5-VL across model design, data construction, and training at various stages, hoping that this report can inspire further research. Seed1.5-VL is now accessible at https://www.volcengine.com/ (Volcano Engine Model ID: doubao-1-5-thinking-vision-pro-250428)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。