arXiv:2607.02674cs.CV2026-07

用自然语言精准生成3D人脸表情,让虚拟形象更生动

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions

论文配图:EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions
图 1 · 摘自论文原文
  • 将表情生成转化为解耦参数空间的回归问题,提升控制精度
  • 构建首个细粒度文本标注的3D表情数据集Txt2Emote,含显性和隐性描述
  • 基于多模态大模型实现高保真3D表情合成,适合虚拟人与动画应用

从文本精确控制3D人脸表情对虚拟化身、动画和人机交互至关重要,但现有方法联合生成身份、表情和纹理,难以实现细粒度控制。本文将文本驱动的表情合成建模为3D可变形模型(3DMM)解耦参数空间中的回归问题。为此,我们引入Txt2Emote数据集,利用GPT-4o与高保真面部追踪器获取多样化3D表情的细粒度文本标注,包含显性特征描述与情境隐含描述。基于此数据集,提出EmoteGPT框架,采用多模态大语言模型(MLLM)并引入专用<Expr>标记语义锚定表情表示,再解码为3DMM参数。通过大规模图像到3DMM数据增强训练,该方法在情感识别指标与感知表达力上超越现有最佳文本到3D人脸生成方法。集成至化身流程后,支持从文本生成逼真或风格化3D化身,以及3D一致的2D人脸表情合成。

原文摘要 · Abstract (English)

Precise control of 3D facial expressions from text is crucial for virtual avatars, animation, and human-computer interaction, yet existing text-to-3D methods jointly generate identity, expression, and texture, making fine-grained expression control difficult. We instead formulate text-driven expression synthesis as a regression problem in the disentangled parameter space of a 3D Morphable Model (3DMM). This setting, however, requires paired data linking detailed language to precise expression parameters, which are missing from existing resources. To fill this gap, we introduce Txt2Emote, a benchmark of diverse 3D facial expressions with fine-grained textual annotations obtained from GPT-4o and a high-fidelity face tracker, providing both explicit descriptions detailing facial features and implicit descriptions referencing the situational context behind the expression. Leveraging this dataset, we present EmoteGPT, a text-to-3D expression framework based on a Multimodal Large Language Model (MLLM) with a dedicated <Expr> token to semantically ground expression representations, which are then decoded into 3DMM parameters. We further improve EmoteGPT by augmenting training with large-scale image-to-3DMM data, enabling it to surpass state-of-the-art text-to-3D face synthesis methods on emotion recognition metrics and in perceived expressiveness. Integrated into avatar pipelines, our method enables photorealistic and stylized 3D avatars, as well as expressive 3D-consistent 2D face synthesis from textual input.

3D人脸文本生成虚拟形象多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。