让图像生成过程像人一样动态思考,提升创意与效率。
Show, Don't Tell: Morphing Latent Reasoning into Image Generation

- 在连续潜在空间中隐式推理,避免文本解码的瓶颈。
- 比基线模型提升16%~25%,推理时间减少44%。
- 适合追求高效、自然生成的视觉创作研究者。
文本到图像(T2I)生成已取得显著进展,但现有方法往往缺乏生成过程中的动态推理与自我修正能力——这是人类创造力的标志。当前增强推理的范式多依赖显式思维过程,在固定步骤将中间推理解码为离散文本,频繁进行图像解码与重编码,导致效率低下、信息丢失和认知错配。为此,我们提出LatentMorph,一种将隐式潜在推理无缝融入T2I生成的新框架。其核心包含四个轻量组件:(i) 编码器,将中间生成状态压缩为紧凑视觉记忆;(ii) 翻译器,将潜在想法转化为可操作指导;(iii) 形塑器,动态引导下一步图像标记预测;(iv) 强化学习训练的触发器,自适应决定推理时机。通过在连续潜在空间中执行推理,LatentMorph避免了显式推理的瓶颈,实现更灵活的自修正。大量实验表明,LatentMorph在基础模型Janus-Pro上,于GenEval提升16%,T2I-CompBench提升25%;在抽象推理任务WISE和IPV-Txt上,优于显式范式(如TwiG)15%和11%;同时推理时间降低44%,标记消耗减少51%;在推理触发的认知对齐度上达到71%,接近人类直觉。
原文摘要 · Abstract (English)
Text-to-image (T2I) generation has achieved remarkable progress, yet existing methods often lack the ability to dynamically reason and refine during generation--a hallmark of human creativity. Current reasoning-augmented paradigms most rely on explicit thought processes, where intermediate reasoning is decoded into discrete text at fixed steps with frequent image decoding and re-encoding, leading to inefficiencies, information loss, and cognitive mismatches. To bridge this gap, we introduce LatentMorph, a novel framework that seamlessly integrates implicit latent reasoning into the T2I generation process. At its core, LatentMorph introduces four lightweight components: (i) a condenser for summarizing intermediate generation states into compact visual memory, (ii) a translator for converting latent thoughts into actionable guidance, (iii) a shaper for dynamically steering next image token predictions, and (iv) an RL-trained invoker for adaptively determining when to invoke reasoning. By performing reasoning entirely in continuous latent spaces, LatentMorph avoids the bottlenecks of explicit reasoning and enables more adaptive self-refinement. Extensive experiments demonstrate that LatentMorph (I) enhances the base model Janus-Pro by $16\%$ on GenEval and $25\%$ on T2I-CompBench; (II) outperforms explicit paradigms (e.g., TwiG) by $15\%$ and $11\%$ on abstract reasoning tasks like WISE and IPV-Txt, (III) while reducing inference time by $44\%$ and token consumption by $51\%$; and (IV) exhibits $71\%$ cognitive alignment with human intuition on reasoning invocation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。