arXiv:2604.09850cs.CV2026-04

无需训练,让图像生成更平衡地关注前景与背景

Training-Free Object-Background Compositional T2I via Dynamic Spatial Guidance and Multi-Path Pruning

  • 通过动态空间引导机制调节前后景注意力
  • 多路径剪枝保留符合构图约束的生成轨迹
  • 适用于需要精准控制物体与背景关系的场景

现有文本到图像扩散模型虽擅长主体生成,但普遍存在前景偏好,将背景视为被动且未充分优化的副产品。这种失衡影响整体场景连贯性,限制构图控制能力。为此,我们提出一种无需训练的框架,重构扩散采样过程以显式建模前景与背景的交互。方法包含两部分:首先,动态空间引导引入时间步依赖的软门控机制,在扩散过程中调节前景与背景注意力,实现空间平衡生成;其次,多路径剪枝通过多路径潜在空间探索,并利用内部注意力统计与外部语义对齐信号动态过滤候选轨迹,保留更满足物体-背景约束的路径。我们还构建了一个专门用于评估物体-背景构图能力的基准。在多个扩散模型骨干上进行的广泛评估表明,该方法在背景连贯性和物体-背景构图一致性方面均有持续提升。

原文摘要 · Abstract (English)

Existing text-to-image diffusion models, while excelling at subject synthesis, exhibit a persistent foreground bias that treats the background as a passive and under-optimized byproduct. This imbalance compromises global scene coherence and constrains compositional control. To address the limitation, we propose a training-free framework that restructures diffusion sampling to explicitly account for foreground-background interactions. Our approach consists of two key components. First, Dynamic Spatial Guidance introduces a soft, time step dependent gating mechanism that modulates foreground and background attention during the diffusion process, enabling spatially balanced generation. Second, Multi-Path Pruning performs multi-path latent exploration and dynamically filters candidate trajectories using both internal attention statistics and external semantic alignment signals, retaining trajectories that better satisfy object-background constraints. We further develop a benchmark specifically designed to evaluate object-background compositionality. Extensive evaluations across multiple diffusion backbones demonstrate consistent improvements in background coherence and object-background compositional alignment.

文本生成图像扩散模型构图控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。