arXiv:2601.14602cs.CV2026-01

用3D空间当草稿纸,让图像生成更精准可控

3D Space as a Scratchpad for Editable Text-to-Image Generation

论文配图:3D Space as a Scratchpad for Editable Text-to-Image Generation
图 1 · 摘自论文原文
  • 将文本提示分解为3D物体,可自由编辑位置朝向
  • 在GenAI-Bench上文本对齐率提升32%
  • 适合需要精确布局控制的创意设计场景

大型语言模型的进步表明,将中间思考外化为显式工作区(如思维链或工具增强推理)能提升推理能力。然而视觉语言模型缺乏类似的时空推理机制,限制了其生成准确反映几何关系、物体身份和组合意图的图像能力。本文提出空间草稿板概念——一种3D推理基底,连接语言意图与图像合成。给定文本提示后,框架解析主体与背景元素,将其实例化为可编辑的3D网格,并通过智能体场景规划确定摆放位置、朝向和视角。最终将3D布局渲染回图像域,保留物体身份线索,使视觉语言模型生成空间一致且视觉连贯的输出。相比以往基于2D布局的方法,本方法支持直观的3D编辑并可靠传递至最终图像。实验证明,在GenAI-Bench上文本对齐率提升32%,证实显式3D推理对精确可控图像生成的益处。结果揭示了视觉语言模型不仅可在语言层面推敲,也可在空间中推演的新范式。

原文摘要 · Abstract (English)

Recent progress in large language models (LLMs) has shown that reasoning improves when intermediate thoughts are externalized into explicit workspaces, such as chain-of-thought traces or tool-augmented reasoning. Yet, visual language models (VLMs) lack an analogous mechanism for spatial reasoning, limiting their ability to generate images that accurately reflect geometric relations, object identities, and compositional intent. We introduce the concept of a spatial scratchpad -- a 3D reasoning substrate that bridges linguistic intent and image synthesis. Given a text prompt, our framework parses subjects and background elements, instantiates them as editable 3D meshes, and employs agentic scene planning for placement, orientation, and viewpoint selection. The resulting 3D arrangement is rendered back into the image domain with identity-preserving cues, enabling the VLM to generate spatially consistent and visually coherent outputs. Unlike prior 2D layout-based methods, our approach supports intuitive 3D edits that propagate reliably into final images. Empirically, it achieves a 32% improvement in text alignment on GenAI-Bench, demonstrating the benefit of explicit 3D reasoning for precise, controllable image generation. Our results highlight a new paradigm for vision-language models that deliberate not only in language, but also in space. Code and visualizations at https://oindrilasaha.github.io/3DScratchpad/

3D生成视觉推理可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。