通过闭环验证推理,让文本生成图像更准确复杂。
Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning

- 将视觉语言规划与扩散生成深度耦合,每步生成都自动验证。
- 仅用4次采样步数就达到商用模型水平,推理速度大幅提升。
- 适合需要高精度复杂图像生成的研究者和开发者使用。
尽管进展迅速,当前文本到图像(T2I)模型仍主要依赖单步生成范式,难以处理复杂语义,且参数规模增长带来的收益递减。现有多步推理方法受限于无实证的规划幻觉、单一事后反思、长上下文优化不稳及高昂推理延迟。为此,我们提出闭环视觉推理(CLVR)框架,深度融合视觉-语言逻辑规划与像素级扩散生成。CLVR引入步级视觉验证的自动化数据引擎,合成可靠推理轨迹;提出代理提示强化学习(PPRL),通过提炼交错的多模态历史生成显式奖励信号,解决长上下文优化不稳问题。此外,为缓解迭代去噪带来的严重延迟,提出Δ-空间权重合并(DSWM),理论上将对齐权重与现成蒸馏先验融合,使每步推理成本降至仅4次非线性函数评估(NFEs),无需昂贵重蒸馏。大量实验表明,CLVR在多个基准上超越现有开源基线,逼近专有商业模型性能,实现了复杂视觉生成的通用测试时扩展能力。
原文摘要 · Abstract (English)
Despite rapid advancements, current text-to-image (T2I) models predominantly rely on a single-step generation paradigm, which struggles with complex semantics and faces diminishing returns from parameter scaling. While recent multi-step reasoning approaches show promise, they are hindered by ungrounded planning hallucinations lacking verification, monolithic post-hoc reflection, long-context optimization instabilities, and prohibitive inference latency. To overcome these bottlenecks, we propose the Closed-Loop Visual Reasoning (CLVR) framework, a comprehensive system that deeply couples visual-language logical planning with pixel-level diffusion generation. CLVR introduces an automated data engine with step-level visual verification to synthesize reliable reasoning trajectories, and proposes Proxy Prompt Reinforcement Learning (PPRL) to resolve long-context optimization instabilities by distilling interleaved multimodal histories into explicit reward signals for accurate causal attribution. Furthermore, to mitigate the severe latency bottleneck caused by iterative denoising, we propose $Δ$-Space Weight Merge (DSWM), a theoretically grounded method that fuses alignment weights with off-the-shelf distillation priors, reducing the per-step inference cost to just 4 NFEs without requiring expensive re-distillation. Extensive experiments demonstrate that CLVR outperforms existing open-source baselines across multiple benchmarks and approaches the performance of proprietary commercial models, unlocking general test-time scaling capabilities for complex visual generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。