arXiv:2603.23924cs.CV2026-03被引 1

无需训练,通过深度仲裁实现更真实的遮挡图像生成

DepthArb: Training-Free Depth-Arbitrated Generation for Occlusion-Robust Image Synthesis

  • 在去噪过程中动态调节注意力,根据深度关系抑制背景干扰
  • 在OcclBench上遮挡相关指标显著优于基线方法
  • 适合需要精确遮挡关系的图像合成任务

文本到图像模型在多个物体重叠区域常出现遮挡关系错误。现有无训练布局引导方法虽约束2D空间位置,但未显式处理深度相关的注意力竞争,导致概念混淆和不合理遮挡。为此,我们提出DepthArb,一种无需训练的框架,将遮挡生成建模为统一去噪轨迹中的注意力仲裁。核心机制包括:注意力仲裁调制(根据相对深度抑制前景内背景注意力),以及空间紧凑性控制(限制注意力扩散以保持物体连贯性)。由于干扰随生成过程变化,遮挡冲突估计构建共享空间冲突场,自适应加权两项目标。DepthArb通过统一的空间-文本注意力接口,作用于U-Net交叉注意力和MMDiT联合注意力的图像-文本部分,无需模型重训。我们进一步提出OcclBench基准,包含连续相对深度设定和专用遮挡评估指标。在OcclBench及公开基准上的实验表明,DepthArb在多个布局与遮挡指标上超越基线,同时保持良好的文本-图像对齐能力。

原文摘要 · Abstract (English)

Text-to-image models often struggle to synthesize correct occlusion relationships among multiple objects, especially in densely overlapping regions. Many training-free layout-guided methods enforce 2D spatial constraints but do not explicitly resolve depth-dependent attention competition, which can cause concept mixing and implausible occlusion. To address this problem, we propose DepthArb, a training-free framework that formulates occlusion generation as attention arbitration within a unified denoising trajectory. DepthArb employs two core occlusion-control mechanisms: Attention Arbitration Modulation suppresses background-object attention within foreground support according to relative depth, while Spatial Compactness Control limits attention dispersion to preserve object coherence. Because interference varies during generation, Occlusion Conflict Estimation constructs a shared spatial conflict field to adaptively weight both objectives. Through a unified spatial-text attention interface, DepthArb operates on U-Net cross-attention and the image-to-text component of MMDiT joint attention without model retraining. We further introduce OcclBench, a benchmark with continuous relative-depth specifications and occlusion-specific evaluation metrics. Experiments on OcclBench and public benchmarks show that DepthArb improves several layout and occlusion metrics over the evaluated baselines while maintaining competitive text-image alignment.

图像生成遮挡关系注意力机制无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。