用视觉推理增强文本语义,让图像生成更符合意图。
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
- 先理解控制图的隐含语义,再生成图像
- 在多个基准上提升图像质量和语义一致性
- 适合需要精准控制生成内容的研究者
可控图像生成领域虽有显著进展,但现有方法仍难以弥合输入文本提示与目标图像之间的语义鸿沟,常过度依赖低层控制信号推断区域细节。为此,我们提出ControlThinker框架,采用“理解-生成”范式:首先通过激励多模态大模型(MLLM)的视觉推理能力,从控制图中挖掘潜在语义以丰富文本提示;该增强后的语义理解可无缝融入图像生成过程,无需额外复杂修改。为应对控制图带来的不确定性,我们鼓励更广泛的推理路径探索,并利用基于指标的输出奖励模型(ORM)选择最优路径。大量实验表明,ControlThinker有效缩小了原始文本提示与目标图像间的语义差距,在多个基准测试中均显著提升图像质量与语义一致性。代码与模型已开源。
原文摘要 · Abstract (English)
The field of controllable image generation has seen significant advancements, with various architectures improving generation layout consistency with control signals. However, contemporary methods still face challenges in bridging the semantic gap between input text prompts with sparse semantics and the target images, often over-relying on low-level control signals to infer regional details. To address this challenge, we propose ControlThinker, a novel framework that employs a "comprehend-then-generate" paradigm. Firstly, by incentivizing the visual reasoning capability of a MLLM, latent semantics from control images are mined to enrich text prompts. This enriched semantic understanding then seamlessly aids in image generation without the need for additional complex modifications. To further tackle the uncertainty arising from the ambiguity of control images, we encourage broader exploration of reasoning trajectories and select the optimal one using a metric-based output reward model (ORM). Extensive experimental results demonstrate that ControlThinker effectively mitigates the semantic gap between raw text prompts and target images, resulting in improved visual quality and semantic consistency across a wide range of benchmarks. The code and models are available at https://github.com/Maplebb/ControlThinker.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。