arXiv:2511.02580cs.CVcs.AI2025-11中稿 · CVPR被引 2

无需训练即可分层生成图像,保持各层一致性。

TAUE: Training-free Noise Transplant and Cultivation Diffusion Model

  • 将去噪过程中的结构信息注入初始噪声,保持空间连贯性。
  • 通过跨层注意力共享语义线索,提升多层视觉一致性。
  • 支持布局感知编辑、多对象组合等新应用,适合创意设计场景。

尽管文本到图像的扩散模型取得了显著进展,但其仅输出单一扁平化图像的特性,成为需要分层控制的专业应用的关键瓶颈。现有方法或依赖大规模不可获取的数据微调,或虽无需训练但仅能生成孤立前景元素,无法生成完整连贯的场景。为此,我们提出无需训练的噪声移植与培育扩散模型(TAUE),一种无需微调和额外数据的分层图像生成框架。TAUE将中间去噪潜在表示中的全局结构信息嵌入初始噪声,以保持空间连贯性,并通过跨层注意力共享语义线索,确保各层间的上下文与视觉一致性。大量实验表明,TAUE在无需训练的方法中达到领先性能,图像质量接近微调模型,同时显著提升层间一致性。此外,该模型支持布局感知编辑、多对象组合和背景替换等新应用,展现出在真实创作流程中实现交互式分层生成系统的潜力。

原文摘要 · Abstract (English)

Despite the remarkable success of text-to-image diffusion models, their output of a single, flattened image remains a critical bottleneck for professional applications requiring layer-wise control. Existing solutions either rely on fine-tuning with large, inaccessible datasets or are training-free yet limited to generating isolated foreground elements, failing to produce a complete and coherent scene. To address this, we introduce the Training-free Noise Transplantation and Cultivation Diffusion Model (TAUE), a novel framework for layer-wise image generation that requires neither fine-tuning nor additional data. TAUE embeds global structural information from intermediate denoising latents into the initial noise to preserve spatial coherence, and integrates semantic cues through cross-layer attention sharing to maintain contextual and visual consistency across layers. Extensive experiments demonstrate that TAUE achieves state-of-the-art performance among training-free methods, delivering image quality comparable to fine-tuned models while improving inter-layer consistency. Moreover, it enables new applications, such as layout-aware editing, multi-object composition, and background replacement, indicating potential for interactive, layer-separated generation systems in real-world creative workflows.

扩散模型分层生成无需训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。