StructDiff让单图生成更保结构、可控位置,还能灵活调整物体大小和细节。
StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation

- 用自适应感受野保持图像全局与局部分布一致。
- 引入3D位置编码实现对物体位置、尺度和细节的灵活控制。
- 首个在单图生成中使用位置编码进行空间操控,适合图像编辑与创作场景。
本文提出 StructDiff,一种基于单尺度扩散模型的单图像生成框架。单图像生成旨在通过捕捉源图像内部统计特性,合成视觉内容相似但多样化的样本,无需依赖外部数据。然而,现有方法在保留结构布局方面表现不佳,尤其针对具有大刚性物体或严格空间约束的图像。同时,多数方法缺乏空间可控性,难以引导生成内容的结构或位置。为此,StructDiff引入自适应感受野模块,以维持全局与局部分布的一致性。在此基础上,结合3D位置编码(PE)作为空间先验,实现对生成物体位置、尺度及局部细节的灵活控制。据我们所知,这是首次在单图像生成中探索基于位置编码的空间操控。此外,提出一种基于大语言模型(LLMs)的新评估准则,克服现有客观指标的局限性和用户研究的高成本问题。StructDiff在文本引导图像生成、图像编辑、外延生成和画图转图像等下游任务中展现出广泛适用性。大量实验表明,其在结构一致性、视觉质量和空间可控性上均优于现有方法。
原文摘要 · Abstract (English)
This paper introduces StructDiff, a generative framework based on a single-scale diffusion model for single-image generation. Single-image generation aims to synthesize diverse samples with similar visual content to the source image by capturing its internal statistics, without relying on external data. However, existing methods often struggle to preserve the structural layout, especially for images with large rigid objects or strict spatial constraints. Moreover, most approaches lack spatial controllability, making it difficult to guide the structure or placement of generated content. To address these challenges, StructDiff introduces an \textit{adaptive receptive field} module to maintain both global and local distributions. Building on this foundation, StructDiff incorporates 3D positional encoding (PE) as a spatial prior, allowing flexible control over positions, scale, and local details of generated objects. To our knowledge, this spatial control capability represents the first exploration of PE-based manipulation in single-image generation. Furthermore, we propose a novel evaluation criterion for single-image generation based on large language models (LLMs). This criterion specifically addresses the limitations of existing objective metrics and the high labor costs associated with user studies. StructDiff also demonstrates broad applicability across downstream tasks, such as text-guided image generation, image editing, outpainting, and paint-to-image synthesis. Extensive experiments demonstrate that StructDiff outperforms existing methods in structural consistency, visual quality, and spatial controllability. The project page is available at https://butter-crab.github.io/StructDiff/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。