提出ControlRef框架,实现高效精准的多实例图像生成。
ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE

- 用统一控制掩码分离实例间语义交互,确保区域绑定精确。
- 引入锚定4D-RoPE,保留空间先验,稀疏布局下推理延迟降80%。
- 适合需要高精度布局控制的多实例图像生成任务。
布局引导的多实例生成对多模态扩散变压器(MM-DiTs)中的可控图像合成至关重要。然而,将此能力整合到统一架构中仍具挑战。现有方法依赖冗余全分辨率画布填充和移位RoPE来管理多参考图像,导致稀疏布局下计算开销剧增,并破坏关键低频RoPE特征,造成严重空间-频率权衡,模糊绝对空间对应关系。为此,我们提出ControlRef,一种高效且精确的多实例合成框架。ControlRef采用统一实例-布局控制(UILC)注意力掩码,严格解耦实例间语义交互并强制精确区域绑定。为进一步促进区域级空间对齐,引入锚定4D-RoPE,一种新型位置编码机制,直接将标记锚定至其绝对几何中心。通过预对齐参考图像至对应边界框分辨率,物理锚定布局与参考标记至其绝对几何中心,并沿z轴堆叠参考图像,锚定4D-RoPE原生保留空间先验,无需有损移位即可缓解空间-频率权衡。大量实验表明,ControlRef在视觉保真度和定位精度上达到当前最佳水平,同时在稀疏布局中推理延迟降低超过80%,在密集场景中内存开销减少50%。
原文摘要 · Abstract (English)
Layout-guided multi-instance generation is essential for controllable image synthesis in Multi-Modal Diffusion Transformers (MM-DiTs). However, integrating this capability into unified architectures remains challenging. Prior frameworks rely on redundant full-resolution canvas padding and Shifted-RoPE to manage multiple reference images. This mechanism drastically inflates computational overhead for sparse layouts and disrupts critical low-frequency RoPE features, creating a severe spatial-frequency compromise that blurs absolute spatial correspondence. To overcome these limitations, we propose ControlRef, a highly efficient and precise multi-instance synthesis framework. ControlRef utilizes a Unified Instance-Layout Control (UILC) attention mask to strictly decouple inter-instance semantic interactions and enforce precise regional binding. To further promote region-level spatial alignment, we introduce Anchored 4D-RoPE, a novel positional encoding mechanism that directly anchors tokens to their absolute geometric centers. By pre-aligning reference images to their corresponding bounding box resolutions, physically anchoring both layout and reference tokens to their absolute geometric centers, and stacking the references along the z-axis, Anchored 4D-RoPE natively preserves spatial priors and mitigates the spatial-frequency compromise without lossy shifting. Extensive experiments demonstrate that ControlRef achieves state-of-the-art visual fidelity and localization accuracy, while concurrently slashing inference latency by over 80% in sparse layouts and reducing memory overhead by 50% in dense scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。