arXiv:2606.24849cs.CVcs.AI2026-06被引 2

让AI画画时更懂布局和数量关系,提升精准度。

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

论文配图:IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
图 1 · 摘自论文原文
  • 将图像生成拆成结构规划和外观渲染两步,分阶段处理
  • 在不依赖草图的情况下,训练时用草图监督提升结构理解
  • 适合需要精确布局的文本生成图像任务

统一的多模态大模型在文本到图像生成上已表现出色,但仍难以准确遵循包含物体数量、空间关系、属性绑定和粗略布局的结构化提示。我们发现这一问题部分源于结构规划与外观渲染在单一条件流中纠缠。为此,提出隐式视觉思维链(IV-CoT),一种用于查询条件图像生成的潜在视觉推理框架。IV-CoT将视觉条件查询分解为结构到语义的级联过程:先生成潜在视觉计划,再基于该计划渲染外观。为引导结构查询,引入仅在训练时使用的草图监督,使模型能从草图中学习结构,而无需在推理时提取草图或中间解码。IV-CoT在单次前向传播中完成隐式思维链推理,在GenEval和T2I-CompBench上表现更优。可视化与分析表明,所学的结构与语义查询在结构感知生成中发挥互补作用。

原文摘要 · Abstract (English)

Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt following, where object counts, spatial relations, attribute bindings, and coarse layouts must be preserved. We attribute this limitation in part to the entanglement of structural planning and appearance rendering within a single conditioning stream. To address this issue, we propose Implicit Visual Chain-of-Thought (IV-CoT), a latent visual reasoning framework for query-conditioned image generation. IV-CoT decomposes the visual conditioning queries into a structural-to-semantic cascade, where structural queries first form a latent visual plan and semantic queries then render appearance conditioned on this plan. To guide the structural queries, we introduce training-only sketch supervision, which encourages them to capture structure from sketches without requiring sketch extraction or intermediate decoding at inference time. IV-CoT performs implicit CoT reasoning in a single forward pass and achieves superior results on GenEval and T2I-CompBench. Visualizations and analyses demonstrate that the learned structural and semantic queries play complementary roles in structure-aware generation.

图像生成结构感知视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。