开源多模态模型InternVL3.5,推理更强、速度更快,支持图形界面与智能体交互。
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- 分阶段强化学习提升逻辑推理能力,先离线后在线优化。
- 推理性能提升16.0%,推理速度加快4.05倍,保持高精度。
- 支持图形界面操作和智能体任务,适合研究与应用开发。
我们推出InternVL 3.5,一个在通用性、推理能力和推理效率上显著超越前代的开源多模态模型系列。核心创新是分层强化学习(Cascade RL)框架,通过离线强化学习实现稳定收敛,再通过在线强化学习进行精细对齐,显著提升下游推理任务表现,如MMMU和MathVista。为优化效率,提出视觉分辨率路由器(ViR),动态调整视觉令牌分辨率而不损失性能;结合解耦视觉-语言部署(DvD)策略,将视觉编码器与语言模型分置于不同GPU,有效平衡计算负载。这些改进使InternVL3.5相较前代InternVL3实现最高16.0%的推理性能提升和4.05×的推理加速。此外,该系列支持图形界面交互与具身智能体能力。其最大模型InternVL3.5-241B-A28B在开源多模态大模型中达到顶尖水平,全面领先于现有开源模型,在通用多模态、推理、文本及智能体任务上接近商业领先模型GPT-5的表现。所有模型与代码均已公开。
原文摘要 · Abstract (English)
We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the Cascade Reinforcement Learning (Cascade RL) framework, which enhances reasoning through a two-stage process: offline RL for stable convergence and online RL for refined alignment. This coarse-to-fine training strategy leads to substantial improvements on downstream reasoning tasks, e.g., MMMU and MathVista. To optimize efficiency, we propose a Visual Resolution Router (ViR) that dynamically adjusts the resolution of visual tokens without compromising performance. Coupled with ViR, our Decoupled Vision-Language Deployment (DvD) strategy separates the vision encoder and language model across different GPUs, effectively balancing computational load. These contributions collectively enable InternVL3.5 to achieve up to a +16.0\% gain in overall reasoning performance and a 4.05$\times$ inference speedup compared to its predecessor, i.e., InternVL3. In addition, InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency. Notably, our largest model, i.e., InternVL3.5-241B-A28B, attains state-of-the-art results among open-source MLLMs across general multimodal, reasoning, text, and agentic tasks -- narrowing the performance gap with leading commercial models like GPT-5. All models and code are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。