用语义编码器实现高分辨率视觉统一理解与生成
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
- 基于多模态大模型和对比语义编码器提取图像特征
- 仅用270万数据即在理解、生成、操控任务上表现优异
- 适合关注视觉生成与感知的开发者和研究者
现有统一模型在视觉语言理解与文本到图像生成方面表现强劲,但在图像感知与操作能力上仍受限,而实际应用中此类能力日益重要。OpenAI推出的GPT-4o-Image模型展示了强大的综合图像感知与操作能力,引发广泛关注。通过实验发现,该模型可能依赖语义编码器而非传统VAE进行特征提取,尽管后者常被视为图像操作的关键。受此启发,我们提出UniWorld-V1,一种基于多模态大语言模型和对比语义编码器构建的统一生成框架。仅使用270万训练数据,该框架在图像理解、生成、操作和感知等多样化任务中均取得优异表现。我们已完全开源UniWorld-V1,包括模型权重、训练与评估脚本及数据集,以促进可复现性与后续研究。
原文摘要 · Abstract (English)
Although existing unified models achieve strong performance in vision-language understanding and text-to-image generation, they remain limited in addressing image perception and manipulation -- capabilities increasingly demanded in practical applications. Recently, OpenAI introduced the powerful GPT-4o-Image model, which showcases advanced capabilities in comprehensive image perception and manipulation, sparking widespread interest. Through carefully designed experiments, we observe that GPT-4o-Image likely relies on semantic encoders rather than VAEs for feature extraction, despite VAEs being commonly regarded as crucial for image manipulation tasks. Inspired by this insight, we propose UniWorld-V1, a unified generative framework built upon semantic features extracted from powerful multimodal large language models and contrastive semantic encoders. Using only 2.7M training data, UniWorld-V1 achieves impressive performance across diverse tasks, including image understanding, generation, manipulation, and perception. We fully open-source the UniWorld-V1 framework, including model weights, training and evaluation scripts, and datasets to promote reproducibility and further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。