arXiv:2501.01998cs.CVcs.AI2025-01IJCAI被引 1

提升Stable Diffusion的3D空间布局能力,让生成图像更符合真实空间关系。

SmartSpatial: Enhancing the 3D Spatial Arrangement Capabilities of Stable Diffusion Models and Introducing a Novel 3D Spatial Evaluation Framework

  • 通过深度信息注入与注意力引导,精准控制物体空间位置。
  • 在空间准确率上显著优于现有方法,实现更高空间保真度。
  • 提供兼顾计算评估与艺术审美的新型3D空间评价框架,适合创意设计场景。

Stable Diffusion模型在文本生成逼真图像方面取得显著进展,但在复杂3D空间关系的准确表达上仍表现不足。为此,我们提出SmartSpatial,一种增强Stable Diffusion空间布局能力的新方法,通过3D感知条件输入与注意力引导机制,结合深度信息注入和交叉注意力控制,实现物体位置的精确放置,显著提升空间准确率指标。同时,我们构建了SmartSpatialEval,一个融合计算空间精度与定性艺术评估的综合性评价框架。实验表明,SmartSpatial在多项空间保真度任务中超越现有方法,为AI驱动的艺术创作树立新基准。

原文摘要 · Abstract (English)

Stable Diffusion models have made remarkable strides in generating photorealistic images from text prompts but often falter when tasked with accurately representing complex spatial arrangements, particularly involving intricate 3D relationships. To address this limitation, we introduce SmartSpatial, an innovative approach that not only enhances the spatial arrangement capabilities of Stable Diffusion but also fosters AI-assisted creative workflows through 3D-aware conditioning and attention-guided mechanisms. SmartSpatial incorporates depth information injection and cross-attention control to ensure precise object placement, delivering notable improvements in spatial accuracy metrics. In conjunction with SmartSpatial, we present SmartSpatialEval, a comprehensive evaluation framework that bridges computational spatial accuracy with qualitative artistic assessments. Experimental results show that SmartSpatial significantly outperforms existing methods, setting new benchmarks for spatial fidelity in AI-driven art and creativity.

3D生成空间布局扩散模型图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。