用强化学习让现有模型学会图文交错生成,无需大量标注数据。
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
- 基于强化学习的后训练策略,融合文本与图像生成的联合优化。
- 在MMIE和InterleavedBench上显著提升图文交错输出的质量与连贯性。
- 适合需要多轮视觉推理或视觉故事生成的研究者使用。
统一视觉语言模型在多模态理解与生成方面取得显著进展,但在生成图文交错输出方面仍存在明显不足,而这种能力对视觉叙事和逐步视觉推理等任务至关重要。本文提出一种基于强化学习的后训练策略,使现有统一模型具备此能力,且无需依赖大规模多模态交错数据集。首先通过混合数据集进行预热训练,包含精心筛选的交错序列及有限的多模态理解与文生图数据,使模型接触交错生成模式的同时保留预训练能力。为进一步优化交错生成,提出统一的策略优化框架,将组相对策略优化(GRPO)扩展至多模态场景。该方法在单一解码轨迹中联合建模文本与图像生成,并采用新型混合奖励函数,涵盖文本相关性、视觉-文本对齐度与结构保真度。此外,引入过程级奖励以提供步骤级指导,提升复杂多模态任务的训练效率。在MMIE与InterleavedBench上的实验表明,本方法显著提升了多模态交错生成的质量与连贯性。
原文摘要 · Abstract (English)
Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, which is a crucial capability for tasks like visual storytelling and step-by-step visual reasoning. In this work, we propose a reinforcement learning-based post-training strategy to unlock this capability in existing unified models, without relying on large-scale multimodal interleaved datasets. We begin with a warm-up stage using a hybrid dataset comprising curated interleaved sequences and limited data for multimodal understanding and text-to-image generation, which exposes the model to interleaved generation patterns while preserving its pretrained capabilities. To further refine interleaved generation, we propose a unified policy optimization framework that extends Group Relative Policy Optimization (GRPO) to the multimodal setting. Our approach jointly models text and image generation within a single decoding trajectory and optimizes it with our novel hybrid rewards covering textual relevance, visual-text alignment, and structural fidelity. Additionally, we incorporate process-level rewards to provide step-wise guidance, enhancing training efficiency in complex multimodal tasks. Experiments on MMIE and InterleavedBench demonstrate that our approach significantly enhances the quality and coherence of multimodal interleaved generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。