arXiv:2601.08179cs.CVcs.AI2026-01被引 2

用文字指令生成3D人脸表情变化,实现自然过渡。

Instruction-Driven 3D Facial Expression Generation and Transition

  • 通过文本指令驱动,分解表情与描述的关联关系。
  • 在CK+和CelebV-HQ数据集上优于现有方法。
  • 适合虚拟角色、动画与情感交互系统开发者。

3D虚拟人通常只有六种基本面部表情。为模拟真实情感变化,需实现任意两种表情间的自然过渡。本文提出一种指令驱动的3D人脸表情生成框架,从一张人脸图像出发,根据文本指令将一种表情转换为另一种。引入指令驱动的表情分解模块(IFED),促进多模态学习,捕捉文本描述与表情特征之间的关联。随后提出I2FET方法,结合IFED与顶点重建损失,优化潜在向量的语义理解,生成符合指令的连续表情序列。最后构建表情过渡模型,实现平滑转换。大量实验表明,该模型在CK+和CelebV-HQ数据集上优于当前最优方法。结果证明,本框架可依据文本指令生成表情轨迹,拓展了表情表达的多样性与应用潜力。

原文摘要 · Abstract (English)

A 3D avatar typically has one of six cardinal facial expressions. To simulate realistic emotional variation, we should be able to render a facial transition between two arbitrary expressions. This study presents a new framework for instruction-driven facial expression generation that produces a 3D face and, starting from an image of the face, transforms the facial expression from one designated facial expression to another. The Instruction-driven Facial Expression Decomposer (IFED) module is introduced to facilitate multimodal data learning and capture the correlation between textual descriptions and facial expression features. Subsequently, we propose the Instruction to Facial Expression Transition (I2FET) method, which leverages IFED and a vertex reconstruction loss function to refine the semantic comprehension of latent vectors, thus generating a facial expression sequence according to the given instruction. Lastly, we present the Facial Expression Transition model to generate smooth transitions between facial expressions. Extensive evaluation suggests that the proposed model outperforms state-of-the-art methods on the CK+ and CelebV-HQ datasets. The results show that our framework can generate facial expression trajectories according to text instruction. Considering that text prompts allow us to make diverse descriptions of human emotional states, the repertoire of facial expressions and the transitions between them can be expanded greatly. We expect our framework to find various practical applications More information about our project can be found at https://vohoanganh.github.io/tg3dfet/

3D生成表情合成文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。