arXiv:2409.13346cs.CVcs.AI2024-09被引 31

无需微调即可个性化生成,保持身份一致且图像质量高。

Imagine yourself: Tuning-Free Personalized Image Generation

  • 用合成数据和并行注意力结构提升多样性与文本对齐
  • 通过分阶段微调,实现高质量图像生成
  • 适合需要快速个性化生成的创作者与开发者

扩散模型在多种图像到图像任务中表现卓越。本文提出Imagine yourself,一种先进的无微调个性化图像生成模型。与传统依赖微调的方法不同,该模型采用共享框架,无需用户单独调整。此前方法在保留身份特征、遵循复杂提示及保持图像质量之间难以平衡,常出现强烈复制粘贴效应,导致生成图像变化有限,如面部表情、姿态改变能力弱,多样性低。为解决这些问题,本文提出:1)新的合成成对数据生成机制以增强图像多样性;2)三文本编码器的全并行注意力架构与可训练视觉编码器,提升文本忠实度;3)粗到细的多阶段微调策略,逐步提升视觉质量。实验表明,Imagine yourself在身份保留、视觉质量和文本对齐方面均优于当前最优模型。人类评估验证其在所有维度上均具领先优势。

原文摘要 · Abstract (English)

Diffusion models have demonstrated remarkable efficacy across various image-to-image tasks. In this research, we introduce Imagine yourself, a state-of-the-art model designed for personalized image generation. Unlike conventional tuning-based personalization techniques, Imagine yourself operates as a tuning-free model, enabling all users to leverage a shared framework without individualized adjustments. Moreover, previous work met challenges balancing identity preservation, following complex prompts and preserving good visual quality, resulting in models having strong copy-paste effect of the reference images. Thus, they can hardly generate images following prompts that require significant changes to the reference image, \eg, changing facial expression, head and body poses, and the diversity of the generated images is low. To address these limitations, our proposed method introduces 1) a new synthetic paired data generation mechanism to encourage image diversity, 2) a fully parallel attention architecture with three text encoders and a fully trainable vision encoder to improve the text faithfulness, and 3) a novel coarse-to-fine multi-stage finetuning methodology that gradually pushes the boundary of visual quality. Our study demonstrates that Imagine yourself surpasses the state-of-the-art personalization model, exhibiting superior capabilities in identity preservation, visual quality, and text alignment. This model establishes a robust foundation for various personalization applications. Human evaluation results validate the model's SOTA superiority across all aspects (identity preservation, text faithfulness, and visual appeal) compared to the previous personalization models.

个性化生成扩散模型无微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。