用扩散模型指导自回归模型,让其更精准生成指定主体图像。
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
- 用扩散模型作为教师,微调自回归模型以提升主体生成能力。
- 微调后自回归模型在主体保真度和提示符合度上超越原扩散模型。
- 特别适合多主体组合与上下文理解任务,展现弱到强的泛化潜力。
基于自回归(AR)架构的多模态模型在文本到图像(T2I)生成等任务中表现出色。然而,我们的研究发现,这类模型在主体驱动的图像生成方面仍不及主流扩散模型。为此,我们提出Proxy-Tuning方法,利用扩散模型来增强自回归模型在特定主体生成上的能力。实验揭示了一种引人注目的弱到强现象:经过微调的自回归模型在主体保真度和提示遵循度上始终优于其扩散模型教师。我们分析了这一性能跃迁,并识别出自回归模型在多主体构图与上下文理解方面的优势场景。该工作不仅在主体驱动的自回归图像生成上取得显著成果,还揭示了图像生成领域中‘弱到强’泛化的潜在可能,深化了对不同架构优劣的理解。
原文摘要 · Abstract (English)
Multimodal autoregressive (AR) models, based on next-token prediction and transformer architecture, have demonstrated remarkable capabilities in various multimodal tasks including text-to-image (T2I) generation. Despite their strong performance in general T2I tasks, our research reveals that these models initially struggle with subject-driven image generation compared to dominant diffusion models. To address this limitation, we introduce Proxy-Tuning, leveraging diffusion models to enhance AR models' capabilities in subject-specific image generation. Our method reveals a striking weak-to-strong phenomenon: fine-tuned AR models consistently outperform their diffusion model supervisors in both subject fidelity and prompt adherence. We analyze this performance shift and identify scenarios where AR models excel, particularly in multi-subject compositions and contextual understanding. This work not only demonstrates impressive results in subject-driven AR image generation, but also unveils the potential of weak-to-strong generalization in the image generation domain, contributing to a deeper understanding of different architectures' strengths and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。