让扩散模型像人一样分步思考,提升图像生成准确性。
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents

- 在扩散模型中引入递归稀疏专家机制,分步优化视觉特征。
- 在ImageNet上实现更高保真度图像生成,优于基线模型。
- 适合需要精细控制与结构化理解的图像生成任务研究者。
扩散模型在高质量数据合成方面取得成功,但在复杂结构化推理(如文本跟随)任务上的能力仍受限。尽管语言模型已通过潜在空间推理和递归机制增强文本理解,但将此类方法应用于多模态文生图任务面临挑战,因视觉标记具有连续且非离散特性。为此,我们借鉴模块化人类认知,提出一种集成于传统扩散模型的递归稀疏专家框架。该方法在联合注意力层中引入递归组件,通过多层潜空间迭代精炼视觉标记,同时利用稀疏选择神经模块实现参数高效共享。每一步由门控网络动态选择特定神经模块,条件依赖于当前视觉标记、扩散时间步及条件信息。在类别条件图像生成任务的ImageNet上进行综合评估,并在GenEval与DPG基准上开展额外研究,结果表明所提方法显著提升了图像生成性能。
原文摘要 · Abstract (English)
Diffusion models have achieved success in high-fidelity data synthesis, yet their capacity for more complex, structured reasoning like text following tasks remains constrained. While advances in language models have leveraged strategies such as latent reasoning and recursion to enhance text understanding capabilities, extending these to multimodal text-to-image generation tasks is challenging due to the continuous and non-discrete nature of visual tokens. To tackle this problem, we draw inspiration from modular human cognition and propose a recursive, sparse mixture-of-experts framework integrated into conventional diffusion models. Our approach introduces a recursive component within joint attention layers that iteratively refines visual tokens over multiple latent steps while efficiently sharing parameters via sparse selection of neural modules. At each step, a gating network is devised to dynamically select specialized neural modules, conditioned on the current visual tokens, the diffusion timestep, and the conditioning information. Comprehensive evaluation on class-conditioned ImageNet image generation tasks and additional studies on the GenEval and DPG benchmark demonstrate the superiority of the proposed method in enhancing model image generation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。