无需训练即可实现风格一致的快速图像生成
A Training-Free Style-aligned Image Generation with Scale-wise Autoregressive Model
- 分尺度自回归模型结合特征替换与动态风格注入
- 风格对齐度显著提升,推理速度超基线6倍以上
- 适合需要快速生成且风格统一的应用场景
我们提出一种无需训练的风格对齐图像生成方法,采用分尺度自回归模型。尽管大规模文本到图像(T2I)模型,特别是基于扩散的方法,在生成质量上表现优异,但常出现生成图像集间风格不一致及推理速度慢的问题,限制了实际应用。为解决这些问题,我们设计三个关键组件:初始特征替换以保证背景一致性,关键特征插值以对齐物体位置,以及动态风格注入,通过调度函数强化风格一致性。与需微调或额外训练的方法不同,本方法保持快速推理的同时保留内容细节。大量实验表明,该方法生成质量可媲美现有方案,显著提升风格对齐度,推理速度超过最快基线模型六倍以上。
原文摘要 · Abstract (English)
We present a training-free style-aligned image generation method that leverages a scale-wise autoregressive model. While large-scale text-to-image (T2I) models, particularly diffusion-based methods, have demonstrated impressive generation quality, they often suffer from style misalignment across generated image sets and slow inference speeds, limiting their practical usability. To address these issues, we propose three key components: initial feature replacement to ensure consistent background appearance, pivotal feature interpolation to align object placement, and dynamic style injection, which reinforces style consistency using a schedule function. Unlike previous methods requiring fine-tuning or additional training, our approach maintains fast inference while preserving individual content details. Extensive experiments show that our method achieves generation quality comparable to competing approaches, significantly improves style alignment, and delivers inference speeds over six times faster than the fastest model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。