让AI更精准理解提示词细节,生成图像更符合要求。
FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation
- 分步拆解提示词,逐项检查生成图像是否达标
- 通过局部优化,使复杂场景生成准确率显著提升
- 适合需要精细控制图像细节的研究与应用
随着多模态大模型的快速发展,统一型多模态大语言模型在图像理解与生成任务中已取得显著进展。然而,尽管这类模型具备自我反思与自优化的能力,其在文本到图像生成中的应用仍较少被探索。现有基于多模态推理的生成方法多依赖提示词增强或整体图文对齐判断,缺乏对提示词具体属性的细粒度反思与优化,导致控制能力有限。为此,我们提出FiRe:一种基于细粒度多模态推理的图像生成方法。FiRe通过将提示词分解为关键视觉要素,先自判断生成图像中各要素的满足程度,再根据自生成的精确反馈进行局部优化。此外,为强化模型推理能力,我们设计了针对FiRe的强化学习方法FiRe-GRPO。由于标准组相对策略优化(GRPO)在多步推理中面临稀疏、结果导向的奖励问题,我们将推理过程建模为步骤级决策问题,设计步骤级奖励,并计算步骤级优势,实现更精细的信用分配。大量实验表明,FiRe在多个基准测试上持续优于主流文本到图像基线,尤其在组合性文本到图像任务中表现突出。
原文摘要 · Abstract (English)
With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced significantly. However, despite the inherent reasoning capabilities of unified MLLMs for self-reflection and self-refinement, their use in text-to-image generation remains largely underexplored. Meanwhile, existing multimodal reasoning-based image generation methods mostly rely on prompt augmentation or holistic image-text alignment judgments, without fine-grained reflection and refinement of detailed prompt attributes, leading to limited fine-grained control. To address this limitation, we propose FiRe, a Fine-grained Multimodal Reasoning method for enhanced image generation by MLLM. In specific, FiRe performs a fine-grained multi-step reasoning by first decomposing the prompt into key visual requirements and then self-judging their satisfaction in the generated image, followed by localized refinement according to self-generated precise feedback. In addition, to further strengthen the MLLM's multimodal reasoning ability, we introduce FiRe-GRPO, a reinforcement learning method tailored to FiRe. Since standard Group Relative Policy Optimization (GRPO) suffers from sparse, outcome-based rewards in multi-step reasoning, we formulate our reasoning process as a step-level decision-making problem, design step-specific rewards, and compute step-level advantages for granular credit assignment within GRPO. Extensive experiments demonstrate that FiRe consistently outperforms competitive text-to-image baselines, including existing reasoning-based methods, with particularly substantial gains on compositional text-to-image benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。