无需训练,通过逐步优化实现高保真物体级图像生成。
COLLAR: Cascaded Object-Level Latent Refinement for High-Fidelity Conditional Generation

- 分阶段扩展视场,用注意力机制注入物体特征
- 在COCO-MIG/POS上优于现有方法,提升语义对齐与空间精度
- 适合需要精细物体控制的生成任务
尽管引入了深度图、Canny边缘等结构先验,扩散变换器在实现高保真物体级控制方面仍面临挑战。现有方法常出现视觉伪影,难以在小范围区域内精确控制物体。为此,我们提出无需训练的级联物体级潜在特征优化框架COLLAR,通过视场(FoV)扩展逐步优化物体级特征。首先,提出跨尺度语义对齐(CSSA)模块,利用注意力机制将物体特征注入扩展视场分支以弥合空间-语义差异;为进一步优化特征,设计循环特征注入(CFI)模块,引入背景反馈机制,采用基于频率的自适应策略选择性地将上下文对齐的局部信息更新至全局主干网络。最终,扩展视场分支作为特征优化枢纽,确保物体特征融入全局生成过程而不降低图像质量。在COCO-MIG和COCO-POS基准上的大量实验表明,本方法在语义对齐、图像质量和空间保真度方面持续超越当前最优方法。
原文摘要 · Abstract (English)
Achieving high-fidelity object-level control in Diffusion Transformers remains a significant challenge despite the introduction of structural priors like depth and Canny maps. Current object-level conditional generation methods frequently suffer from visual artifacts and struggle to maintain precise control over objects within small localized regions. To address these limitations, we propose Cascaded Object-Level Latent Refinement (COLLAR), a training-free framework that progressively optimizes object-level features via the Field-of-View (FoV) expansion. First, we propose the Cross-Scale Semantic Alignment (CSSA) module to address spatial-semantic gaps by injecting object-level features into extended-FoV branches via attention mechanisms. To further optimize these features, the Cyclic Feature Injection (CFI) module introduces a reciprocal background feedback mechanism. It leverages a frequency-based adaptive strategy to selectively update the global backbone with context-aligned local information. Finally, the extended-FoV branch serves as a hub for feature optimization, ensuring that object-level features are integrated into the global generation process without compromising final image quality. Extensive experiments on the COCO-MIG and COCO-POS benchmarks demonstrate that our approach consistently outperforms state-of-the-art methods across semantic alignment, image quality, and spatial fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。