用自学习表情特征实现无需地标约束的3D人脸动画迁移
FreeAvatar: Robust 3D Facial Animation Transfer by Learning an Expression Foundation Model
- 构建表情基础模型,从无标签数据中学习细粒度表情表征
- 通过可微神经渲染实现端到端动画迁移,保持表情一致性
- 单网络支持多角色联合训练,适配任意野外输入人脸
视频驱动的3D人脸动画迁移旨在让虚拟形象复现演员的表情。现有方法虽在几何与感知一致性上取得进展,但基于面部关键点的几何约束难以捕捉细微情绪,而分类任务训练的表达特征缺乏复杂情绪的精细粒度。为此,我们提出FreeAvatar,一种仅依赖自学习表达表征的鲁棒人脸动画迁移方法。其由表达基础模型与人脸动画迁移模型构成:前者通过人脸重建任务构建特征空间,并利用大量未标注面部图像及新收集的表情对比数据集优化表达特征空间;后者提出表达驱动的多角色动画器,将表达语义映射至3D角色的控制参数,并通过输入输出图像间的感知约束维持表情一致性。整个流程通过训练好的神经渲染器实现可微转换。此外,不同于以往需为每个角色单独解码的方法,我们设计动态身份注入模块,使多个角色可在单一网络中联合训练。
原文摘要 · Abstract (English)
Video-driven 3D facial animation transfer aims to drive avatars to reproduce the expressions of actors. Existing methods have achieved remarkable results by constraining both geometric and perceptual consistency. However, geometric constraints (like those designed on facial landmarks) are insufficient to capture subtle emotions, while expression features trained on classification tasks lack fine granularity for complex emotions. To address this, we propose \textbf{FreeAvatar}, a robust facial animation transfer method that relies solely on our learned expression representation. Specifically, FreeAvatar consists of two main components: the expression foundation model and the facial animation transfer model. In the first component, we initially construct a facial feature space through a face reconstruction task and then optimize the expression feature space by exploring the similarities among different expressions. Benefiting from training on the amounts of unlabeled facial images and re-collected expression comparison dataset, our model adapts freely and effectively to any in-the-wild input facial images. In the facial animation transfer component, we propose a novel Expression-driven Multi-avatar Animator, which first maps expressive semantics to the facial control parameters of 3D avatars and then imposes perceptual constraints between the input and output images to maintain expression consistency. To make the entire process differentiable, we employ a trained neural renderer to translate rig parameters into corresponding images. Furthermore, unlike previous methods that require separate decoders for each avatar, we propose a dynamic identity injection module that allows for the joint training of multiple avatars within a single network.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。