轻量控制模块让视觉自回归模型高效实现复杂图像生成调控
Efficient Conditional Generation on Scale-based Visual Autoregressive Models
- 用轻量控制模块+分布式架构,实时融合条件信号
- 训练成本降低,生成质量超越现有方法
- 适合需要快速部署可控图像生成的开发者
近期自回归(AR)模型在图像生成上已可媲美扩散模型。但复杂空间条件生成仍依赖微调,成本高昂。本文提出高效控制模型(ECM),采用轻量控制模块与分布式架构,通过上下文感知注意力层实时优化条件特征,并共享门控前馈网络以提升容量利用率和控制特征一致性。针对早期生成对语义结构的关键作用,引入早期优先采样策略,减少每轮训练令牌数以降低计算开销;推理时辅以温度调度补偿后期令牌训练不足。在基于尺度的自回归模型上的大量实验表明,该方法在保持高保真度和多样性的同时,显著提升了训练与推理效率,优于现有基线。
原文摘要 · Abstract (English)
Recent advances in autoregressive (AR) models have demonstrated their potential to rival diffusion models in image synthesis. However, for complex spatially-conditioned generation, current AR approaches rely on fine-tuning the pre-trained model, leading to significant training costs. In this paper, we propose the Efficient Control Model (ECM), a plug-and-play framework featuring a lightweight control module that introduces control signals via a distributed architecture. This architecture consists of context-aware attention layers that refine conditional features using real-time generated tokens, and a shared gated feed-forward network (FFN) designed to maximize the utilization of its limited capacity and ensure coherent control feature learning. Furthermore, recognizing the critical role of early-stage generation in determining semantic structure, we introduce an early-centric sampling strategy that prioritizes learning early control sequences. This approach reduces computational cost by lowering the number of training tokens per iteration, while a complementary temperature scheduling during inference compensates for the resulting insufficient training of late-stage tokens. Extensive experiments on scale-based AR models validate that our method achieves high-fidelity and diverse control over image generation, surpassing existing baselines while significantly improving both training and inference efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。