arXiv:2603.16936cs.CVcs.AI2026-03

用语言模型统一建模面部动作,实现文本驱动动画与理解

TDMM-LM: Bridging Facial Understanding and Animation via Language Models

  • 构建80小时合成面部视频数据集,配以3D参数与文本描述
  • 语言模型可准确描述面部动态,并根据文本生成对应动作
  • 首次将面部参数建模为语言问题,适合动画与人机交互研究者

文本引导的人体动作动画已快速发展,但面部动画因缺乏高质量的标注文本对数据而滞后。为此,我们利用基础生成模型合成大规模、均衡的面部行为数据集。设计覆盖情绪与头部动作的提示模板,通过多个生成器生成约80小时面部视频,并拟合每帧3D面部参数,形成大规模(提示与参数)对用于训练。基于该数据集,我们通过两个互补任务探查语言模型在面部动作上的双向能力:(1) Motion2Language:给定3D面部参数序列,模型生成包含内容、风格与动态的自然语言描述;(2) Language2Motion:给定提示,模型通过量化动作标记合成对应的3D面部参数序列以支持下游动画。大量实验表明,语言模型在此设置下具备强泛化能力,能同时解析与生成面部动作。据我们所知,这是首个将面部参数建模视为语言问题的工作,为文本条件下的面部动画与动作理解建立了统一路径。

原文摘要 · Abstract (English)

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced corpus of facial behavior. We design prompts suite covering emotions and head motions, generate about 80 hours of facial videos with multiple generators, and fit per-frame 3D facial parameters, yielding large-scale (prompt and parameter) pairs for training. Building on this dataset, we probe language models for bidirectional competence over facial motion via two complementary tasks: (1) Motion2Language: given a sequence of 3D facial parameters, the model produces natural-language descriptions capturing content, style, and dynamics; and (2) Language2Motion: given a prompt, the model synthesizes the corresponding sequence of 3D facial parameters via quantized motion tokens for downstream animation. Extensive experiments show that in this setting language models can both interpret and synthesize facial motion with strong generalization. To best of our knowledge, this is the first work to cast facial-parameter modeling as a language problem, establishing a unified path for text-conditioned facial animation and motion understanding.

面部动画语言模型3D参数文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。