arXiv:2412.01485cs.CV2024-12CVPR被引 3

先标准化再个性化,生成人物图像更像且更可控。

SerialGen: Personalized Image Generation by First Standardization Then Personalization

  • 分两步生成:先统一参考图风格,再根据提示词生成
  • 生成图像与参考图整体外观一致,对文本提示响应准确
  • 适合需要连续一致人物形象的创作场景

本文关注个性化人物图像生成中高文本控制力与全身外观一致性问题。提出名为SerialGen的新框架,采用两阶段串行生成方法:第一阶段为标准化阶段,统一参考图像;第二阶段基于标准化后的参考图进行个性化生成。此外,引入两个模块以增强标准化过程。实验表明,该框架能生成忠实还原参考图全身外观、并精准响应多样文本提示的个性化图像。深入分析证实,串行生成方法与标准化模型对提升参考图与输出图间、以及不同文本提示下序列输出间的外观一致性具有关键作用。'Serial'一词兼具双重含义:既指两阶段方法,也强调生成序列图像时保持外观一致的能力。

原文摘要 · Abstract (English)

In this work, we are interested in achieving both high text controllability and whole-body appearance consistency in the generation of personalized human characters. We propose a novel framework, named SerialGen, which is a serial generation method consisting of two stages: first, a standardization stage that standardizes reference images, and then a personalized generation stage based on the standardized reference. Furthermore, we introduce two modules aimed at enhancing the standardization process. Our experimental results validate the proposed framework's ability to produce personalized images that faithfully recover the reference image's whole-body appearance while accurately responding to a wide range of text prompts. Through thorough analysis, we highlight the critical contribution of the proposed serial generation method and standardization model, evidencing enhancements in appearance consistency between reference and output images and across serial outputs generated from diverse text prompts. The term "Serial" in this work carries a double meaning: it refers to the two-stage method and also underlines our ability to generate serial images with consistent appearance throughout.

图像生成个性化一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。