7B小模型实现顶级文本渲染,单卡即可部署。
Ovis-Image Technical Report
- 用强多模态主干+专注文本的训练方案
- 性能媲美大模型,接近闭源系统
- 单张高端显卡就能运行,适合实际应用
我们提出Ovis-Image,一个70亿参数的文本到图像模型,专为高质量文本渲染优化,可在严苛计算约束下高效运行。基于之前的Ovis-U1框架,Ovis-Image融合基于扩散的视觉解码器与更强的Ovis 2.5多模态主干,采用以文本为中心的训练流程,结合大规模预训练与精细化后训练调整。尽管模型规模紧凑,其文本渲染性能可与显著更大的开源模型如Qwen-Image比肩,并逼近闭源系统如Seedream和GPT4o。关键在于,该模型仍可在单张高端GPU上以中等显存完成部署,大幅缩小前沿文本渲染能力与实际可用性之间的差距。结果表明,结合强大多模态主干与精心设计的文本导向训练策略,足以实现可靠的双语文本渲染,无需依赖超大规模或专有模型。
原文摘要 · Abstract (English)
We introduce $\textbf{Ovis-Image}$, a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational constraints. Built upon our previous Ovis-U1 framework, Ovis-Image integrates a diffusion-based visual decoder with the stronger Ovis 2.5 multimodal backbone, leveraging a text-centric training pipeline that combines large-scale pre-training with carefully tailored post-training refinements. Despite its compact architecture, Ovis-Image achieves text rendering performance on par with significantly larger open models such as Qwen-Image and approaches closed-source systems like Seedream and GPT4o. Crucially, the model remains deployable on a single high-end GPU with moderate memory, narrowing the gap between frontier-level text rendering and practical deployment. Our results indicate that combining a strong multimodal backbone with a carefully designed, text-focused training recipe is sufficient to achieve reliable bilingual text rendering without resorting to oversized or proprietary models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。