无需训练,用现成专家模型精准控制人脸生成细节。
ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
- 直接调用预训练的人脸识别、属性识别等专家模型进行引导。
- 在扩散过程中保持图像真实性和分布一致性,实现高精度控制。
- 支持多专家协作,可同时调节身份、年龄、属性等多种特征。
近期扩散模型在文本到人脸生成方面取得显著进展,但对人脸特征的细粒度控制仍具挑战。现有方法通常需额外训练模块来处理身份、属性或年龄等特定控制,灵活性差且资源消耗大。本文提出ExpertGen,一种无需训练的框架,利用预训练的专家模型(如人脸识别、面部属性识别、年龄估计网络)对生成过程进行精细引导。方法采用潜在一致性模型,在每个扩散步骤中确保生成结果真实且符合分布,从而提供有效引导信号以精确控制生成过程。实验表明,专家模型可高精度引导生成,多个专家可协同工作,实现对多种面部特征的同步控制。通过直接集成现成专家模型,该方法使任意此类模型均可作为即插即用组件用于可控人脸生成。
原文摘要 · Abstract (English)
Recent advances in diffusion models have significantly improved text-to-face generation, but achieving fine-grained control over facial features remains a challenge. Existing methods often require training additional modules to handle specific controls such as identity, attributes, or age, making them inflexible and resource-intensive. We propose ExpertGen, a training-free framework that leverages pre-trained expert models such as face recognition, facial attribute recognition, and age estimation networks to guide generation with fine control. Our approach uses a latent consistency model to ensure realistic and in-distribution predictions at each diffusion step, enabling accurate guidance signals to effectively steer the diffusion process. We show qualitatively and quantitatively that expert models can guide the generation process with high precision, and multiple experts can collaborate to enable simultaneous control over diverse facial aspects. By allowing direct integration of off-the-shelf expert models, our method transforms any such model into a plug-and-play component for controllable face generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。