通过分阶段调节与双平衡机制,实现高保真人脸定制化生成。
PersonaMagic: Stage-Regulated High-Fidelity Face Customization with Tandem Equilibrium
- 分阶段学习面部概念嵌入,提升生成精度
- 双平衡机制使身份保留与文本描述更协调
- 可通用至非人脸领域,适配现有个性化模型
个性化图像生成在适应新概念方面已取得显著进展,但如何在准确重建未见概念的同时保持根据提示的可编辑性,尤其是在处理面部特征的复杂细节时,仍是挑战。本文研究文本到图像条件生成过程中的时间动态,强调阶段划分对引入新概念的关键作用。提出PersonaMagic,一种分阶段调控的生成方法,用于高保真人脸定制。通过一个简单的MLP网络,在特定时间步区间内学习一系列嵌入以捕捉人脸概念。同时,设计了双平衡机制,调整文本编码器中的自注意力响应,平衡文本描述与身份保留,显著提升两者表现。大量实验表明,PersonaMagic在定性和定量评估上均优于当前最优方法。其鲁棒性与灵活性也在非人脸领域得到验证,还可作为插件增强预训练个性化模型性能。
原文摘要 · Abstract (English)
Personalized image generation has made significant strides in adapting content to novel concepts. However, a persistent challenge remains: balancing the accurate reconstruction of unseen concepts with the need for editability according to the prompt, especially when dealing with the complex nuances of facial features. In this study, we delve into the temporal dynamics of the text-to-image conditioning process, emphasizing the crucial role of stage partitioning in introducing new concepts. We present PersonaMagic, a stage-regulated generative technique designed for high-fidelity face customization. Using a simple MLP network, our method learns a series of embeddings within a specific timestep interval to capture face concepts. Additionally, we develop a Tandem Equilibrium mechanism that adjusts self-attention responses in the text encoder, balancing text description and identity preservation, improving both areas. Extensive experiments confirm the superiority of PersonaMagic over state-of-the-art methods in both qualitative and quantitative evaluations. Moreover, its robustness and flexibility are validated in non-facial domains, and it can also serve as a valuable plug-in for enhancing the performance of pretrained personalization models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。