通过双向协同生成,提升图文空间理解的准确性。
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
- 构建共享的3D场景图,统一表示图文空间关系。
- 在VSD数据集上显著超越主流图文生成方法。
- 适合需要精准空间对齐的视觉语言任务研究者。
在视觉空间理解(VSU)领域,空间图像到文本(SI2T)与空间文本到图像(ST2I)是成对出现的两个基础任务。现有独立的SI2T或ST2I方法在空间理解上表现不佳,主要源于三维空间特征建模困难。本文提出在双学习框架下联合建模这两项任务。为此,我们引入一种新型3D场景图(3DSG)表示,可共享并同时促进两项任务。进一步地,基于3D→图像与3D→文本过程在ST2I与SI2T中具有对称性的直觉,提出空间双离散扩散(SD³)框架,利用3D→X的中间特征引导X→3D的困难过程,使整体生成能力相互增强。在视觉空间理解数据集VSD上,本系统显著优于主流T2I与I2T方法。深入分析揭示了双学习策略的推进机制。
原文摘要 · Abstract (English)
In the visual spatial understanding (VSU) area, spatial image-to-text (SI2T) and spatial text-to-image (ST2I) are two fundamental tasks that appear in dual form. Existing methods for standalone SI2T or ST2I perform imperfectly in spatial understanding, due to the difficulty of 3D-wise spatial feature modeling. In this work, we consider modeling the SI2T and ST2I together under a dual learning framework. During the dual framework, we then propose to represent the 3D spatial scene features with a novel 3D scene graph (3DSG) representation that can be shared and beneficial to both tasks. Further, inspired by the intuition that the easier 3D$\to$image and 3D$\to$text processes also exist symmetrically in the ST2I and SI2T, respectively, we propose the Spatial Dual Discrete Diffusion (SD$^3$) framework, which utilizes the intermediate features of the 3D$\to$X processes to guide the hard X$\to$3D processes, such that the overall ST2I and SI2T will benefit each other. On the visual spatial understanding dataset VSD, our system outperforms the mainstream T2I and I2T methods significantly. Further in-depth analysis reveals how our dual learning strategy advances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。