10B小模型实现顶尖多模态能力,性能超大模型十倍。
STEP3-VL-10B Technical Report
- 统一预训练+全量微调,融合视觉与语言理解。
- 在MMBench达92.2%,超越10-20倍大的模型。
- 适合追求高效推理的开发者与研究者使用。
我们提出STEP3-VL-10B,一个轻量级开源基础模型,旨在重新定义小型化与前沿多模态智能之间的平衡。该模型通过两项关键策略实现:首先,在1.2万亿多模态标记上采用统一且全参数未冻结的预训练策略,结合语言对齐的视觉编码器与Qwen3-8B解码器,建立内在的视觉-语言协同机制;其次,采用扩展后的后训练流程,包含超过1000轮强化学习。关键创新在于引入并行协同推理(PaCoRe),可扩展测试时计算资源,用于探索和合成多样化的视觉假设。尽管仅具100亿参数,其性能仍媲美甚至超越10至20倍更大的模型(如GLM-4.6V-106B、Qwen3-VL-235B)以及顶级专有模型如Gemini 2.5 Pro与Seed-1.5-VL。在多项评测中表现卓越:MMBench达92.2%,MMMU为80.11%,复杂推理任务中AIME2025达94.43%,MathVision为75.95%。我们开放全部模型组件,为社区提供强大、高效且可复现的基础基准。
原文摘要 · Abstract (English)
We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-VL-10B is realized through two strategic shifts: first, a unified, fully unfrozen pre-training strategy on 1.2T multimodal tokens that integrates a language-aligned Perception Encoder with a Qwen3-8B decoder to establish intrinsic vision-language synergy; and second, a scaled post-training pipeline featuring over 1k iterations of reinforcement learning. Crucially, we implement Parallel Coordinated Reasoning (PaCoRe) to scale test-time compute, allocating resources to scalable perceptual reasoning that explores and synthesizes diverse visual hypotheses. Consequently, despite its compact 10B footprint, STEP3-VL-10B rivals or surpasses models 10$\times$-20$\times$ larger (e.g., GLM-4.6V-106B, Qwen3-VL-235B) and top-tier proprietary flagships like Gemini 2.5 Pro and Seed-1.5-VL. Delivering best-in-class performance, it records 92.2% on MMBench and 80.11% on MMMU, while excelling in complex reasoning with 94.43% on AIME2025 and 75.95% on MathVision. We release the full model suite to provide the community with a powerful, efficient, and reproducible baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。