arXiv:2410.04965cs.CV2024-10被引 3

用扩散模型提升3D人脸编辑的精准度与通用性。

Revealing Directions for Text-guided 3D Face Editing

  • 在生成模型隐空间中通过正反文本提示扩散,定位可编辑区域。
  • 支持任意属性描述,实现对3D人脸局部区域的精确控制。
  • 适用于多种预训练GAN,适合多媒体内容创作场景。

3D人脸编辑是多媒体领域的重要任务,旨在通过不同控制信号操纵3D人脸模型。3D感知生成对抗网络(GAN)能仅从单张2D图像学习出表达丰富的3D模型,促使研究者探索其隐空间中的语义编辑方向。然而,现有方法在质量、效率和泛化性之间难以平衡。为此,本文探索将扩散模型的优势引入3D感知GAN。我们提出Face Clan,一种基于任意属性描述快速生成并操控3D人脸的新方法。为实现解耦编辑,我们通过一对相反提示在隐空间进行扩散,估计出目标区域的掩码;随后对掩码区域的隐码执行去噪,揭示编辑方向。该方法提供精确可控的编辑能力,用户可通过文本描述直观指定关注区域。实验表明,Face Clan在多种预训练GAN上均具有效性和泛化能力,为文本引导的人脸编辑提供了直观且广泛应用的解决方案。

原文摘要 · Abstract (English)

3D face editing is a significant task in multimedia, aimed at the manipulation of 3D face models across various control signals. The success of 3D-aware GAN provides expressive 3D models learned from 2D single-view images only, encouraging researchers to discover semantic editing directions in its latent space. However, previous methods face challenges in balancing quality, efficiency, and generalization. To solve the problem, we explore the possibility of introducing the strength of diffusion model into 3D-aware GANs. In this paper, we present Face Clan, a fast and text-general approach for generating and manipulating 3D faces based on arbitrary attribute descriptions. To achieve disentangled editing, we propose to diffuse on the latent space under a pair of opposite prompts to estimate the mask indicating the region of interest on latent codes. Based on the mask, we then apply denoising to the masked latent codes to reveal the editing direction. Our method offers a precisely controllable manipulation method, allowing users to intuitively customize regions of interest with the text description. Experiments demonstrate the effectiveness and generalization of our Face Clan for various pre-trained GANs. It offers an intuitive and wide application for text-guided face editing that contributes to the landscape of multimedia content creation.

3D人脸编辑文本生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。