arXiv:2411.15732cs.GRcs.CV2024-11

用扩散模型生成逼真动态人脸,支持精准提示编辑。

DynamicAvatars: Accurate Dynamic Facial Avatars Reconstruction and Precise Editing with Diffusion Models

  • 基于高斯溅射与双跟踪框架重建动态3D人脸
  • 通过大语言模型生成引导参数,实现精确编辑
  • 适合虚拟现实与影视制作中的高精度人脸建模

生成和编辑动态3D头部形象在虚拟现实和影视制作中至关重要。现有方法常存在面部失真、头部动作不准确及细粒度编辑能力有限的问题。为此,我们提出DynamicAvatars,一种从视频片段和面部位置/表情参数生成逼真动态3D头像的动态模型。该方法通过新颖的基于提示的编辑模型实现精确编辑,将用户提示与由大语言模型(LLMs)生成的引导参数结合。我们提出基于高斯溅射的双跟踪框架,并引入提示预处理模块以提升编辑稳定性。通过集成专用GAN算法并连接控制模块(从LLMs生成精确引导参数),有效克服了现有方法的局限性。此外,我们设计了一种动态编辑策略,选择性使用特定训练数据集,提升了模型在动态编辑任务中的效率与适应性。

原文摘要 · Abstract (English)

Generating and editing dynamic 3D head avatars are crucial tasks in virtual reality and film production. However, existing methods often suffer from facial distortions, inaccurate head movements, and limited fine-grained editing capabilities. To address these challenges, we present DynamicAvatars, a dynamic model that generates photorealistic, moving 3D head avatars from video clips and parameters associated with facial positions and expressions. Our approach enables precise editing through a novel prompt-based editing model, which integrates user-provided prompts with guiding parameters derived from large language models (LLMs). To achieve this, we propose a dual-tracking framework based on Gaussian Splatting and introduce a prompt preprocessing module to enhance editing stability. By incorporating a specialized GAN algorithm and connecting it to our control module, which generates precise guiding parameters from LLMs, we successfully address the limitations of existing methods. Additionally, we develop a dynamic editing strategy that selectively utilizes specific training datasets to improve the efficiency and adaptability of the model for dynamic editing tasks.

3D人脸扩散模型动态生成提示编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。