统一多模态模型UniGen通过测试时验证提升图像生成质量。
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- 引入链式思维验证机制,测试时自动生成并评估图像与文本对齐度。
- 在GenEval上取得0.78分,DPG-Bench达85.19,性能领先。
- 全开源数据训练,适合研究统一多模态模型构建的开发者。
我们提出UniGen,一种能够实现图像理解与生成的统一多模态大语言模型。从数据驱动视角出发,研究了UniGen从多阶段预训练、监督微调到直接偏好优化的完整训练流程。更重要的是,提出一种新的测试时链式思维验证(CoT-V)策略,显著提升图像生成质量。CoT-V使UniGen在测试时既能生成图像,又能以逐步推理方式验证文本与图像的语义一致性。所有阶段均使用开源数据训练,UniGen在多个图像理解与生成基准上达到顶尖表现:GenEval得分为0.78,DPG-Bench达85.19。通过大量消融实验,本工作为构建统一多模态大模型提供了可操作的洞见,指明了未来研究方向。
原文摘要 · Abstract (English)
We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More importantly, we propose a new Chain-of-Thought Verification (CoT-V) strategy for test-time scaling, which significantly boosts UniGen's image generation quality using a simple Best-of-N test-time strategy. Specifically, CoT-V enables UniGen to act as both image generator and verifier at test time, assessing the semantic alignment between a text prompt and its generated image in a step-by-step CoT manner. Trained entirely on open-source datasets across all stages, UniGen achieves state-of-the-art performance on a range of image understanding and generation benchmarks, with a final score of 0.78 on GenEval and 85.19 on DPG-Bench. Through extensive ablation studies, our work provides actionable insights and addresses key challenges in the full life cycle of building unified MLLMs, contributing meaningful directions to the future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。