arXiv:2509.23452cs.CVcs.CL2025-09

让AI理解非镜头视角的空间描述,精准调整图像方向与深度。

FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

  • 引入参考坐标系统一语言与视觉空间表达,实现对齐评估。
  • 单次修正使顶尖文生图模型性能提升最高7.0个百分点。
  • 适合需要精确空间控制的创意设计、虚拟场景生成任务。

当前文本到图像生成模型在非相机视角的空间描述上表现不佳。为此,我们提出基于大模型的扩散编辑框架FoR-SALE,作为自校正大模型控制扩散模型(SLD)的扩展。该方法首先评估文本与初始生成图像的对齐度,再根据空间描述中的参考坐标系(FoR)修正图像。通过视觉模块提取图像空间结构,并将空间表达映射至对应相机视角,实现语言与视觉的统一对齐。检测到偏差后,生成并执行相应的编辑操作。FoR-SALE引入新颖的潜在空间操作,用于调整图像朝向和深度。我们在三个针对参考坐标系的空间理解基准上进行评估:FoR-LMD、FoREST和FoREST-G,后者包含多对象场景、扩展关系类型及自然化提示的诊断子集。实验表明,仅需一次修正,本框架可使最先进文生图模型性能最高提升7.0个百分点。

原文摘要 · Abstract (English)

Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are described from non-camera perspectives. To address this limitation, we propose Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing (FoR-SALE), an extension of the Self-correcting LLM-controlled Diffusion (SLD). FoR-SALE first evaluates the alignment between a given text and an initially generated image, and then refines the image based on the expressed FoR in the spatial description. It employs vision modules to extract the spatial configuration of the generated image and simultaneously maps the spatial expression to a corresponding camera perspective. This unified perspective enables direct evaluation of alignment between language and vision. When misalignment is detected, the required editing operations are generated and applied. FoR-SALE introduces novel latent-space operations to adjust the facing direction and depth of generated images. We evaluate FoR-SALE on three benchmarks designed to assess spatial understanding with FoR: FoR-LMD, FoREST, and FoREST-G, with the latter providing diagnostic subsets covering multi-object scenes, expanded relation types, and naturalized prompts to test the generality of our framework. Our framework improves the performance of SOTA T2I models by up to 7.0 pp using only a single round of correction.

文生图空间理解扩散模型大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。