动态选择图像局部控制信号,提升多条件文生图精度与质量
PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
- 按图像区域动态选控,避免多条件冲突
- 分阶段注入控制信息,结构先稳再细化
- 适合需要精准布局的文生图任务
基于扩散模型的文生图技术在视觉条件控制方面取得进展,但现有类似ControlNet的方法在组合式视觉控制上表现不佳:同时保持多个异构控制信号的语义一致性与高视觉质量时,因采用独立控制分支,在去噪过程中常产生冲突引导,导致结构扭曲和伪影。为此,本文提出PixelPonder,一种统一控制框架,可在单一控制结构下实现多视觉条件的有效控制。具体地,设计了像素级自适应条件选择机制,动态优先选取子区域相关的控制信号,实现精准局部引导且无全局干扰;同时引入时间感知的控制注入策略,根据去噪步骤调节控制影响,逐步从结构保持过渡到纹理优化,充分融合不同类别控制信息,促进更协调的图像生成。大量实验表明,PixelPonder在多个基准数据集上超越已有方法,在空间对齐准确率上显著提升,同时保持高文本语义一致性。
原文摘要 · Abstract (English)
Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing ControlNet-like methods struggle with compositional visual conditioning - simultaneously preserving semantic fidelity across multiple heterogeneous control signals while maintaining high visual quality, where they employ separate control branches that often introduce conflicting guidance during the denoising process, leading to structural distortions and artifacts in generated images. To address this issue, we present PixelPonder, a novel unified control framework, which allows for effective control of multiple visual conditions under a single control structure. Specifically, we design a patch-level adaptive condition selection mechanism that dynamically prioritizes spatially relevant control signals at the sub-region level, enabling precise local guidance without global interference. Additionally, a time-aware control injection scheme is deployed to modulate condition influence according to denoising timesteps, progressively transitioning from structural preservation to texture refinement and fully utilizing the control information from different categories to promote more harmonious image generation. Extensive experiments demonstrate that PixelPonder surpasses previous methods across different benchmark datasets, showing superior improvement in spatial alignment accuracy while maintaining high textual semantic consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。