arXiv:2603.18627cs.AI2026-03

通过闭环搜索与空间引导,提升文本生成图像的精准度。

Agentic Flow Steering and Parallel Rollout Search for Spatially Grounded Text-to-Image Generation

  • 用视觉语言模型实时检查中间结果并动态调整生成方向。
  • 在三个基准上超越FLUX.1-dev,实现当前最佳效果。
  • 提供快速版和高性能版,兼顾速度与精度。

精确的文本到图像(T2I)生成已取得显著进展,但受限于静态文本编码器的关系推理能力不足,以及开环采样中的误差累积。缺乏实时反馈导致微小语义模糊在常微分方程轨迹中逐步放大,偏离空间约束。为此,我们提出AFS-Search(Agentic Flow Steering and Parallel Rollout Search),一个基于FLUX.1-dev的免训练闭环框架。AFS-Search结合免训练的闭环并行回溯搜索与流引导机制,利用视觉语言模型(VLM)作为语义判别器,诊断中间潜在表示,并通过精确的空间定位动态调节速度场。同时,将T2I生成建模为序列决策过程,通过前瞻模拟探索多条生成轨迹,并依据VLM引导的奖励选择最优路径。此外,我们还提供了AFS-Search-Pro以提升性能,以及AFS-Search-Fast以加快生成速度。实验表明,AFS-Search-Pro显著提升原FLUX.1-dev表现,在三个不同基准上达到当前最优水平;而AFS-Search-Fast在保持快速生成的同时也显著改善了质量。

原文摘要 · Abstract (English)

Precise Text-to-Image (T2I) generation has achieved great success but is hindered by the limited relational reasoning of static text encoders and the error accumulation in open-loop sampling. Without real-time feedback, initial semantic ambiguities during the Ordinary Differential Equation trajectory inevitably escalate into stochastic deviations from spatial constraints. To bridge this gap, we introduce AFS-Search (Agentic Flow Steering and Parallel Rollout Search), a training-free closed-loop framework built upon FLUX.1-dev. AFS-Search incorporates a training-free closed-loop parallel rollout search and flow steering mechanism, which leverages a Vision-Language Model (VLM) as a semantic critic to diagnose intermediate latents and dynamically steer the velocity field via precise spatial grounding. Complementarily, we formulate T2I generation as a sequential decision-making process, exploring multiple trajectories through lookahead simulations and selecting the optimal path based on VLM-guided rewards. Further, we provide AFS-Search-Pro for higher performance and AFS-Search-Fast for quicker generation. Experimental results show that our AFS-Search-Pro greatly boosts the performance of the original FLUX.1-dev, achieving state-of-the-art results across three different benchmarks. Meanwhile, AFS-Search-Fast also significantly enhances performance while maintaining fast generation speed.

文本生成图像闭环生成空间引导VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。