用单图生成带夸张表情的3D人脸,支持表情和文字控制。
Joker: Conditional 3D Head Synthesis with Extreme Facial Expressions
- 结合2D扩散模型与3DMM,实现多模态表情控制。
- 在雕塑、画作等非真实图像上仍保持高质量生成效果。
- 首次实现视角一致的极端舌头动作合成,适合影视特效应用。
我们提出Joker,一种基于单张参考图像生成具有极端面部表情的3D人头的新方法。通过3D形态模型(3DMM)与文本输入联合控制表情,解决仅用3DMM难以刻画细微情绪变化及涉及口腔、舌部动作的极端表情问题。该方法基于2D扩散先验,具备良好的域外泛化能力,适用于雕塑、浓妆、绘画等非真实样本,同时保持高表达力。为提升视图一致性,提出一种新型3D蒸馏技术,将2D先验预测转换为神经辐射场(NeRF)。实验表明,该方法在多项指标上达到当前最优,且据我们所知,首次实现了视角一致的极端舌部运动合成。
原文摘要 · Abstract (English)
We introduce Joker, a new method for the conditional synthesis of 3D human heads with extreme expressions. Given a single reference image of a person, we synthesize a volumetric human head with the reference identity and a new expression. We offer control over the expression via a 3D morphable model (3DMM) and textual inputs. This multi-modal conditioning signal is essential since 3DMMs alone fail to define subtle emotional changes and extreme expressions, including those involving the mouth cavity and tongue articulation. Our method is built upon a 2D diffusion-based prior that generalizes well to out-of-domain samples, such as sculptures, heavy makeup, and paintings while achieving high levels of expressiveness. To improve view consistency, we propose a new 3D distillation technique that converts predictions of our 2D prior into a neural radiance field (NeRF). Both the 2D prior and our distillation technique produce state-of-the-art results, which are confirmed by our extensive evaluations. Also, to the best of our knowledge, our method is the first to achieve view-consistent extreme tongue articulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。