arXiv:2506.23044cs.CVcs.AI2025-06被引 65

Ovis-U1是30亿参数的多模态统一模型,支持图文理解、文生图和图像编辑。

Ovis-U1 Technical Report

  • 从语言模型出发的统一训练,融合理解与生成任务提升性能
  • 在多模态基准上得分69.6,文生图评测达83.72分,图像编辑超主流模型
  • 适合需要端到端多模态能力的研究与应用开发者

本文介绍Ovis-U1,一个30亿参数的统一模型,整合了多模态理解、文生图生成和图像编辑功能。基于Ovis系列架构,Ovis-U1采用基于扩散的视觉解码器与双向标记精炼器,使生成能力媲美GPT-4o。不同于以往使用冻结多模态大模型进行生成的做法,Ovis-U1从语言模型出发,采用统一训练范式,在理解与生成任务联合训练下表现更优。在OpenCompass多模态学术基准上取得69.6分,超越Ristretto-3B和SAIL-VL-1.5-2B等近期先进模型。文生图任务中,DPG-Bench得分为83.72,GenEval得分为0.89;图像编辑任务中,ImgEdit-Bench得分为4.00,GEdit-Bench-EN得分为6.42。作为Ovis统一模型系列的首版,Ovis-U1显著推动了多模态理解、生成与编辑的边界。

原文摘要 · Abstract (English)

In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Building on the foundation of the Ovis series, Ovis-U1 incorporates a diffusion-based visual decoder paired with a bidirectional token refiner, enabling image generation tasks comparable to leading models like GPT-4o. Unlike some previous models that use a frozen MLLM for generation tasks, Ovis-U1 utilizes a new unified training approach starting from a language model. Compared to training solely on understanding or generation tasks, unified training yields better performance, demonstrating the enhancement achieved by integrating these two tasks. Ovis-U1 achieves a score of 69.6 on the OpenCompass Multi-modal Academic Benchmark, surpassing recent state-of-the-art models such as Ristretto-3B and SAIL-VL-1.5-2B. In text-to-image generation, it excels with scores of 83.72 and 0.89 on the DPG-Bench and GenEval benchmarks, respectively. For image editing, it achieves 4.00 and 6.42 on the ImgEdit-Bench and GEdit-Bench-EN, respectively. As the initial version of the Ovis unified model series, Ovis-U1 pushes the boundaries of multimodal understanding, generation, and editing.

多模态文生图图像编辑统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。