优化图文生成顺序,提升模型对空间关系和多模态理解的能力。
Reinforcing the Generation Order of Multimodal Masked Diffusion Models

- 引入可学习控制模块,用强化学习动态决定生成顺序。
- 在GenEval上提升4.08%,在VLMEvalKit上提升4.85%。
- 适合需要精准图文对齐与复杂多模态推理的研究者。
扩散语言模型(DLMs)在自然语言生成任务中取得了显著进展。近期研究显示,自适应的词元生成顺序能显著提升数学推理与代码合成性能。本文探讨了文本到图像生成与多模态理解中的生成顺序优化问题。我们发现,与数独等结构化语言任务不同,仅靠模型输出概率不足以确定最优生成序列。为此,我们提出一个通过分组相对策略优化(GRPO)训练的可学习控制模块,以决定生成顺序。实验表明,该方法显著提升了DLMs在文本-图像对齐与多模态理解方面的表现。尤其增强了模型对生成图像中细微空间关系的捕捉能力,并改善了多模态推理与理解任务性能。在面向物体聚焦的文本-图像对齐基准GenEval上,实现4.08%的相对提升;在VLMEvalKit测试中,多模态理解性能提升4.85%,验证了该方法的广泛有效性。
原文摘要 · Abstract (English)
Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstrates that adaptive token generation ordering can significantly improve performance in mathematical reasoning and code synthesis applications. In this work, we investigate the optimization of generation order for both text-to-image synthesis and multimodal understanding. We first establish that, unlike structured problems in language generation such as Sudoku puzzles, model logits alone are insufficient for determining optimal generation sequences in text-to-image generation and multimodal understanding. To address this challenge, we introduce a learnable control module trained via Group Relative Policy Optimization (GRPO) to determine the generation order. Our results demonstrate that learning this control block substantially improves both text-to-image alignment and multimodal understanding in DLMs. In particular, it enhances the model's ability to capture fine-grained spatial relationships in generated images while also strengthening performance on multimodal reasoning and comprehension tasks. We evaluate our framework on GenEval, an object-focused benchmark for text-to-image alignment, where it achieves 4.08% relative improvements. In addition, experiments on VLMEvalKit confirm 4.85% relative improvements in multimodal understanding, highlighting the broad effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。