用结构化视觉信息提升文本生成图像的推理能力。
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
- 用文本构建的结构化视觉表示替代中间图像,实现纯文本推理
- 在T2I-ReasonBench上提升4.61%,WISE上提升4%
- 兼容多种生成器,适合需要强推理的图像生成任务
基于推理的文本到图像(T2I)生成要求模型准确理解复杂提示。现有方法分为两类:(1) 纯文本推理,计算高效但缺乏视觉上下文,常遗漏关键空间与视觉元素;(2) 文本-图像交替推理,利用T2I生成器提供视觉参考,虽增强视觉对齐,但计算开销大,且受限于生成器的表征能力。为此,我们提出StruVis,一种通过结构化视觉思考来增强T2I生成的新框架。该框架不依赖中间图像生成,而是使用基于文本的结构化视觉表示作为推理中间状态,使多模态大模型在纯文本过程中有效“感知”视觉结构,从而释放其在推理型T2I生成中的潜力。作为生成器无关的框架,StruVis可无缝集成于多种T2I生成器,并显著提升其性能。大量实验表明,其在推理型T2I基准上取得显著进步,如在T2I-ReasonBench上提升4.61%,在WISE上提升4%。
原文摘要 · Abstract (English)
Reasoning-based text-to-image (T2I) generation requires models to interpret complex prompts accurately. Existing reasoning frameworks can be broadly categorized into two types: (1) Text-Only Reasoning, which is computationally efficient but lacks access to visual context, often resulting in the omission of critical spatial and visual elements; and (2) Text-Image Interleaved Reasoning, which leverages a T2I generator to provide visual references during the reasoning process. While this approach enhances visual grounding, it incurs substantial computational costs and constrains the reasoning capacity of MLLMs to the representational limitations of the generator. To this end, we propose StruVis, a novel framework that enhances T2I generation through Thinking with Structured Vision. Instead of relying on intermediate image generation, StruVis employs text-based structured visual representations as intermediate reasoning states, thereby enabling the MLLM to effectively "perceive" visual structure within a purely text-based reasoning process. Powered by this, the reasoning potential for T2I generation of the MLLM is unlocked through structured-vision-guided reasoning. Additionally, as a generator-agnostic reasoning framework, our proposed StruVis can be seamlessly integrated with diverse T2I generators and efficiently enhance their performance in reasoning-based T2I generation. Extensive experiments demonstrate that StruVis achieves significant performance improvements on reasoning-based T2I benchmarks, e.g., a 4.61% gain on T2I-ReasonBench and a 4% gain on WISE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。