arXiv:2509.20775cs.CVcs.AI2025-09

零样本提升人像定制生成质量与可控性

CusEnhancer: A Zero-Shot Scene and Controllability Enhancement Method for Photo Customization via ResInversion

  • 通过双反向潜在空间融合实现三流统一生成与重建
  • 在场景多样性、身份保真度上达当前最优,且无需重训控制器
  • 提出新型逆向方法ResInversion,提速129倍,显著降低计算开销

近期基于文本到图像扩散模型的人像合成取得显著进展,但现有方法存在场景退化、控制不足和感知身份不一致等问题。本文提出CustomEnhancer,一种零样本增强框架,用于提升现有身份定制模型性能。该框架利用人脸交换技术与预训练扩散模型,以零样本方式获取额外表示并编码至个性化模型。通过提出的三流融合生成方法,识别并结合两个兼容的反向潜在空间,操纵个性化模型的关键空间,统一生成与重建过程,实现三流生成。该流程还实现了无需训练的全面控制,支持精准个性化生成且无需为每模型重训控制器。针对零文本逆向(NTI)高时间复杂度问题,引入ResInversion——一种通过预扩散机制进行噪声修正的新逆向方法,使逆向时间减少129倍。实验表明,CustomEnhancer在场景多样性、身份保真度、训练自由控制方面均达到当前最优水平,同时验证了ResInversion对NTI的效率优势。代码将在论文录用后公开。

原文摘要 · Abstract (English)

Recently remarkable progress has been made in synthesizing realistic human photos using text-to-image diffusion models. However, current approaches face degraded scenes, insufficient control, and suboptimal perceptual identity. We introduce CustomEnhancer, a novel framework to augment existing identity customization models. CustomEnhancer is a zero-shot enhancement pipeline that leverages face swapping techniques, pretrained diffusion model, to obtain additional representations in a zeroshot manner for encoding into personalized models. Through our proposed triple-flow fused PerGeneration approach, which identifies and combines two compatible counter-directional latent spaces to manipulate a pivotal space of personalized model, we unify the generation and reconstruction processes, realizing generation from three flows. Our pipeline also enables comprehensive training-free control over the generation process of personalized models, offering precise controlled personalization for them and eliminating the need for controller retraining for per-model. Besides, to address the high time complexity of null-text inversion (NTI), we introduce ResInversion, a novel inversion method that performs noise rectification via a pre-diffusion mechanism, reducing the inversion time by 129 times. Experiments demonstrate that CustomEnhancer reach SOTA results at scene diversity, identity fidelity, training-free controls, while also showing the efficiency of our ResInversion over NTI. The code will be made publicly available upon paper acceptance.

图像生成零样本可控生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。