arXiv:2609.07815cs.CVcs.AI2026-09

提出视觉思维框架,让图文生成更可解释可控。

VoT: Vision-of-Thought for Unified Multimodal Representation Alignment

  • 在视觉语言模型与扩散模型间加入离散视觉思维层,生成物体布局等高层视觉计划。
  • 通过闭环训练使思维令牌既语义可读又保留生成所需视觉信息。
  • 提升图文对齐效果,支持可解释、可控的图像生成,适合需要精准控制的场景。

当前文本到图像系统普遍采用“文本编码器加扩散解码器”范式,文本语义直接调控连续潜变量噪声。尽管取得成功,这些方法缺乏显式的、可解释的中间表示,难以有效连接高层语言语义与底层视觉信号。本文提出视觉思维(VoT)框架,在视觉语言模型(VLMs)与扩散变换器(DiTs)之间引入一个离散的视觉思考层。不再将VLM仅视为文本编码器,而是用其作为多模态规划器,生成代表高层视觉计划(如物体和布局)的离散VoT令牌,再进行像素渲染。我们在VLM语义空间中训练专用的VoT分词器,采用闭环目标函数,结合VLM对齐、特征重建和向量量化损失。这些目标使令牌在语义上可被VLM理解,同时保留生成所需的视觉信息。实验表明,VoT提升了语义对齐性,并为可解释与可控生成提供了结构化接口。

原文摘要 · Abstract (English)

Current text-to-image systems typically employ a "text encoder plus diffusion decoder" paradigm, in which text semantics directly modulate continuous latent noise. Despite their success, these methods lack an explicit, interpretable intermediate representation that effectively bridges high-level linguistic semantics and low-level visual signals. In this paper, we propose Vision-of-Thought (VoT), a framework that introduces a discrete visual-thinking layer between vision-language models (VLMs) and diffusion transformers (DiTs). Instead of treating VLMs merely as text encoders, we use them as multimodal planners that generate discrete VoT tokens representing high-level visual plans, such as objects and layouts, before rendering pixels. We train a specialized VoT tokenizer in the VLM semantic space with a closed-loop objective that combines VLM alignment, feature reconstruction, and vector-quantization losses. These objectives make the tokens semantically readable by the VLM while preserving the visual information needed for generation. Experimental results demonstrate that VoT improves semantic alignment and provides a structured interface for interpretable and controllable generation.

视觉思维图文生成可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。