无需训练即可实现结构与外观双重精准控制的文生图方法
RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
- 分离条件特征采样与去噪过程,动态调节注入时机
- 在复杂条件下生成图像,结构对齐度与视觉质量均达领先水平
- 适用于多种模型架构,支持组合式条件生成
文生图扩散模型在文本生成高质量图像方面表现卓越。近期工作通过引入条件图像(如canny边缘图)实现细粒度空间控制。其中,特征注入方法作为无需训练的替代方案受到关注,但常出现结构错位、条件泄漏和视觉伪影,尤其当条件图像偏离自然RGB分布时更为明显。我们分析发现,现有方法忽视了条件特征采样调度这一关键因素,未考虑扩散过程中结构保持与域对齐之间的动态平衡。为此,提出一种灵活的无训练框架,将条件特征采样调度与去噪过程解耦,并系统研究特征注入调度谱,以更好平衡结构对齐与外观质量。进一步引入重启精炼调度提升采样过程,结合外观丰富的提示策略改善视觉效果。实验表明,该方法在复杂多样的条件下均达到当前最优性能。其通用性使其天然支持组合式条件生成,并可即插即用地适配从UNet到现代DiT骨干网络(如FLUX)的多种架构。
原文摘要 · Abstract (English)
Text-to-image (T2I) diffusion models have shown remarkable success in generating high-quality images from text prompts. Recent efforts extend these models to incorporate conditional images (e.g., canny edge) for fine-grained spatial control. Among them, feature injection methods have emerged as a training-free alternative to traditional fine-tuning-based approaches. However, they often suffer from structural misalignment, condition leakage, and visual artifacts, especially when the condition image diverges significantly from natural RGB distributions. Through an analysis of existing methods, we identify a key limitation: the sampling schedule of condition features, previously unexplored, fails to account for the evolving interplay between structure preservation and domain alignment throughout diffusion steps. Inspired by this observation, we propose a flexible training-free framework that decouples the sampling schedule of condition features from the denoising process, and systematically investigate the spectrum of feature injection schedules to achieve a better balance between structural alignment and appearance quality. We further enhance the sampling process by introducing a restart refinement schedule, and improve the visual quality with an appearance-rich prompting strategy. Together, these designs enable training-free controllable generation that is both structure-rich and appearance-rich. Extensive experiments demonstrate that our method achieves state-of-the-art performance under complex and diverse conditions. Owing to its generality, our framework naturally supports compositional conditional generation and generalizes across architectures in a plug-and-play manner, from UNet-based diffusion models to modern DiT backbones such as FLUX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。