让文生图更精准还原物体边界和细小结构。
PixelControl: Fine-Grained Condition Fidelity in Text-to-Image Diffusion

- 在像素空间直接控制,避开潜在空间压缩的失真问题。
- 通过多尺度结构损失与结构感知注入,提升边框和小区域准确性。
- 支持深度、分割、边缘等多种条件,适合需要高精度布局的应用。
可控文生图扩散模型虽能遵循全局空间布局,但在物体边界、细线轮廓及中等/小区域等细粒度结构上仍易出错。这一问题在基于VAE的潜在扩散模型中尤为突出,因空间压缩会削弱高频和低面积条件信号。本文提出PixelControl,一种像素空间可控扩散框架,以实现细粒度条件保真。基于PixelDiT架构,避免潜在瓶颈,引入两项互补设计:一是结构感知控制注入,生成条件结构图并强化敏感区域的控制残差;二是多尺度金字塔周期损失,在多个分辨率下验证生成图像与条件结构的一致性,平衡全局布局与局部边界和细节精度。通过模态专用控制分支与轻量级门控融合,支持深度、分割、边缘及其组合。在深度、分割和边缘控制任务上的实验表明,PixelControl在结构保真度和视觉质量上均优于现有方法,尤其在边界和中/小区域表现显著提升。
原文摘要 · Abstract (English)
Controllable text-to-image diffusion models can often follow the global layout of spatial conditions, yet still violate fine-grained structures such as object boundaries, thin contours, and medium/small conditioned regions. This limitation is especially problematic for VAE-based latent diffusion, where spatial compression can weaken high-frequency and low-area condition signals. We propose PixelControl, a pixel-space controllable diffusion framework for fine-grained condition fidelity. Built on a PixelDiT-style backbone, PixelControl avoids the latent bottleneck and introduces two complementary designs. First, Structure-Aware Control Injection derives a condition structure map and uses it to strengthen injected control residuals around spatially sensitive regions. Second, Multi-Scale Pyramid Cycle Loss verifies generated images against condition-derived structures across multiple resolutions, balancing global layout consistency with local boundary and detail accuracy. PixelControl supports depth, segmentation, edge, and their combinations through modality-specific control branches with lightweight gated fusion. Experiments across depth, segmentation, and edge control show that PixelControl improves structural fidelity and visual quality over existing controllable generation methods, with especially strong gains on boundaries and medium/small conditioned regions. The project page can be found at: https://linxin0.github.io/pixelcontrol_homepage/pixelcontrol-site/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。