arXiv:2604.26341cs.CV2026-04

让统一图像生成模型具备内在3D几何感知能力

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness

论文配图:SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness
图 1 · 摘自论文原文
  • 用并行空间变换器增强多模态大模型的3D建模能力
  • 生成显式深度图并注入扩散模型,提升空间一致性
  • 在文本到图像和图像编辑中表现优异,推理开销低

近期统一图像生成模型通过多模态大模型(MLLM)实现语义理解,借助扩散模型生成图像,但在空间感知任务上仍受限于缺乏内在空间理解与生成过程中的显式几何引导。本文提出SpatialFusion框架,将3D几何感知内化至统一图像生成模型。首先采用混合变压器(MoT)架构,在MLLM中引入并行空间变换器,共享自注意力机制,从丰富语义上下文中学习目标图像的度量深度图。这些显式几何结构通过专用深度适配器注入扩散模型骨干,提供精确的空间约束以实现空间一致的图像生成。通过渐进式两阶段训练策略,SpatialFusion在多个空间感知基准上显著优于领先模型(如GPT-4o),并在文本到图像生成与图像编辑任务中实现泛化性能提升,同时保持可忽略的推理开销。

原文摘要 · Abstract (English)

Recent unified image generation models have achieved remarkable success by employing MLLMs for semantic understanding and diffusion backbones for image generation. However, these models remain fundamentally limited in spatially-aware tasks due to a lack of intrinsic spatial understanding and the absence of explicit geometric guidance during generation. In this paper, we propose SpatialFusion, a novel framework that internalizes 3D geometric awareness into unified image generation models. Specifically, we first employ a Mixture-of-Transformers (MoT) architecture to augment the MLLM with a parallel spatial transformer to enhance 3D geometric modeling capability. By sharing self-attention with the MLLM, the spatial transformer learns to derive metric-depth maps of target images from rich semantic contexts. These explicit geometric scaffolds are then injected into the diffusion backbone through a specialized depth adapter, providing precise spatial constraints for spatially-coherent image generation. Through a progressive two-stage training strategy, SpatialFusion significantly enhances performance on spatially-aware benchmarks, notably outperforming leading models such as GPT-4o. Additionally, it achieves generalized performance gains across both text-to-image generation and image editing scenarios, all while maintaining negligible inference overhead.

3D生成图像生成扩散模型几何感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。