Ovis2.5提升视觉感知与多模态推理,支持原生分辨率图像处理。
Ovis2.5 Technical Report
- 采用原生分辨率视觉变压器,保留图像细节与整体布局。
- 引入反思机制,推理准确率显著提升,支持可选思考模式。
- 2B和9B版本均达开源模型同规模最优,适合本地部署。
我们提出Ovis2.5,是专为原生分辨率视觉感知与强多模态推理设计的下一代模型。其集成原生分辨率视觉变换器,以可变分辨率直接处理图像,避免固定分辨率分块带来的失真,有效保留复杂图表等密集视觉内容的精细细节与全局结构。为强化推理能力,模型通过训练实现超越线性思维链的反思机制,包括自检与修正,并在推理时以可选“思考模式”提供延迟与精度权衡。模型通过五阶段综合训练课程逐步构建能力:从基础视觉与多模态预训练,到大规模指令微调,最终通过DPO与GRPO实现对齐与推理增强。为高效扩展,采用多模态数据打包与混合并行策略,实现显著端到端加速。我们发布两个开源模型:Ovis2.5-9B与Ovis2.5-2B。后者延续‘小模型、大性能’理念,适用于资源受限的本地场景。在OpenCompass多模态排行榜上,Ovis2.5-9B平均得分78.3,大幅超越前代Ovis2-8B,成为参数量低于40B的开源多模态大模型中最佳表现;Ovis2.5-2B得分为73.9,创下同类尺寸模型新纪录。除总分外,其在STEM基准测试、定位与视频任务中表现领先,且在复杂图表分析上达到开源模型同规模最优水平。
原文摘要 · Abstract (English)
We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer that processes images at their native, variable resolutions, avoiding the degradation from fixed-resolution tiling and preserving both fine detail and global layout -- crucial for visually dense content like complex charts. To strengthen reasoning, we train the model to move beyond linear chain-of-thought and perform reflection -- including self-checking and revision. This advanced capability is exposed as an optional "thinking mode" at inference time, allowing users to trade latency for enhanced accuracy on difficult inputs. The model is trained via a comprehensive five-phase curriculum that progressively builds its skills. The process begins with foundational visual and multimodal pretraining, advances through large-scale instruction tuning, and culminates in alignment and reasoning enhancement using DPO and GRPO. To scale these upgrades efficiently, we employ multimodal data packing and hybrid parallelism, yielding a significant end-to-end speedup. We release two open-source models: Ovis2.5-9B and Ovis2.5-2B. The latter continues the "small model, big performance" philosophy of Ovis2, making it ideal for resource-constrained, on-device scenarios. On the OpenCompass multimodal leaderboard, Ovis2.5-9B averages 78.3, marking a substantial improvement over its predecessor, Ovis2-8B, and achieving state-of-the-art results among open-source MLLMs in the sub-40B parameter range; Ovis2.5-2B scores 73.9, establishing SOTA for its size. Beyond aggregate scores, Ovis2.5 achieves leading results on STEM benchmarks, exhibits strong capabilities on grounding and video tasks, and achieves open-source SOTA at its scale for complex chart analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。