arXiv:2506.03147cs.CVcs.AI2025-06被引 46

用语义编码器实现高分辨率视觉统一理解与生成

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

  • 基于多模态大模型和对比语义编码器提取图像特征
  • 仅用270万数据即在理解、生成、操控任务上表现优异
  • 适合关注视觉生成与感知的开发者和研究者

现有统一模型在视觉语言理解与文本到图像生成方面表现强劲,但在图像感知与操作能力上仍受限,而实际应用中此类能力日益重要。OpenAI推出的GPT-4o-Image模型展示了强大的综合图像感知与操作能力,引发广泛关注。通过实验发现,该模型可能依赖语义编码器而非传统VAE进行特征提取,尽管后者常被视为图像操作的关键。受此启发,我们提出UniWorld-V1,一种基于多模态大语言模型和对比语义编码器构建的统一生成框架。仅使用270万训练数据,该框架在图像理解、生成、操作和感知等多样化任务中均取得优异表现。我们已完全开源UniWorld-V1,包括模型权重、训练与评估脚本及数据集,以促进可复现性与后续研究。

原文摘要 · Abstract (English)

Although existing unified models achieve strong performance in vision-language understanding and text-to-image generation, they remain limited in addressing image perception and manipulation -- capabilities increasingly demanded in practical applications. Recently, OpenAI introduced the powerful GPT-4o-Image model, which showcases advanced capabilities in comprehensive image perception and manipulation, sparking widespread interest. Through carefully designed experiments, we observe that GPT-4o-Image likely relies on semantic encoders rather than VAEs for feature extraction, despite VAEs being commonly regarded as crucial for image manipulation tasks. Inspired by this insight, we propose UniWorld-V1, a unified generative framework built upon semantic features extracted from powerful multimodal large language models and contrastive semantic encoders. Using only 2.7M training data, UniWorld-V1 achieves impressive performance across diverse tasks, including image understanding, generation, manipulation, and perception. We fully open-source the UniWorld-V1 framework, including model weights, training and evaluation scripts, and datasets to promote reproducibility and further research.

视觉生成语义编码统一框架多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。