让模型一次生成图文,自动切换不依赖人工干预。
Unified Text-Image Generation with Weakness-Targeted Post-Training
- 后训练阶段用自生成数据,实现文本到图像的自动过渡。
- 在4个基准上提升生成质量,尤其针对弱项优化效果显著。
- 适合想构建端到端多模态生成系统的研究者。
统一多模态生成架构可联合生成文本与图像,是文本到图像(T2I)合成的新兴方向。然而,现有系统多依赖显式模态切换,在生成推理文本后手动转为图像生成,这种分步推理限制了跨模态耦合,难以实现自动化多模态生成。本文探索通过后训练实现完全统一的文本-图像生成,使模型在单次推理中自主从文本推理过渡到视觉合成。我们研究了联合生成对T2I性能的影响及后训练中各模态的重要性。进一步对比不同后训练数据策略,发现针对特定缺陷设计的数据集优于通用图像-标题语料库或基准对齐数据。采用离线、基于奖励加权的后训练方法,结合完全自生成的合成数据,本方法在四个不同T2I基准上均实现多模态图像生成性能提升,验证了双模态奖励加权与策略性数据设计的有效性。
原文摘要 · Abstract (English)
Unified multimodal generation architectures that jointly produce text and images have recently emerged as a promising direction for text-to-image (T2I) synthesis. However, many existing systems rely on explicit modality switching, generating reasoning text before switching manually to image generation. This separate, sequential inference process limits cross-modal coupling and prohibits automatic multimodal generation. This work explores post-training to achieve fully unified text-image generation, where models autonomously transition from textual reasoning to visual synthesis within a single inference process. We examine the impact of joint text-image generation on T2I performance and the relative importance of each modality during post-training. We additionally explore different post-training data strategies, showing that a targeted dataset addressing specific limitations achieves superior results compared to broad image-caption corpora or benchmark-aligned data. Using offline, reward-weighted post-training with fully self-generated synthetic data, our approach enables improvements in multimodal image generation across four diverse T2I benchmarks, demonstrating the effectiveness of reward-weighting both modalities and strategically designed post-training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。