arXiv:2512.11321cs.CV2025-12

用自然语言生成带语义的关键帧,让面部动画更可控、更直观。

KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes

  • 通过语义关键帧替代连续帧,实现可解释的面部动画控制。
  • 在2100个脚本数据上,表情还原度和语义对齐度显著提升。
  • 适合需要精细编辑的影视/游戏角色动画创作者。

面部动画是计算机图形学中数字角色创作的核心。传统制作流程依赖稀疏但语义明确的关键帧来精准控制表情。若能直接从自然语言描述生成此类动画,将极大提升内容创作效率与可及性。然而,现有方法多采用文本到连续帧的范式,直接回归密集的面部运动轨迹,导致高层语义意图与底层运动混杂,缺乏显式的语义控制结构,限制了精确编辑与可解释性。受动画制作中关键帧范式的启发,我们提出KeyframeFace,一种基于可解释关键帧的语义面部动画生成框架。该方法不预测密集运动轨迹,而是以可解释的ARKit面部控制空间中的语义关键帧序列表示动画。一个语言驱动模型利用大语言模型先验,生成与上下文文本描述和情感线索对齐的关键帧。为支持该范式,我们构建了一个多模态数据集,包含2,100个表达脚本、单目视频、每帧ARKit系数及人工标注的语义关键帧。实验表明,引入语义关键帧监督与语言先验,显著提升了表情保真度与语义对齐度,优于不使用面部动作语义的方法。

原文摘要 · Abstract (English)

Facial animation is a core component for creating digital characters in Computer Graphics (CG) industry. A typical production workflow relies on sparse, semantically meaningful keyframes to precisely control facial expressions. Enabling such animation directly from natural-language descriptions could significantly improve content creation efficiency and accessibility. However, most existing methods adopt a text-to-continuous-frames paradigm, directly regressing dense facial motion trajectories from language. This formulation entangles high-level semantic intent with low-level motion, lacks explicit semantic control structure, and limits precise editing and interpretability. Inspired by the keyframe paradigm in animation production, we propose KeyframeFace, a framework for semantic facial animation from language via interpretable keyframes. Instead of predicting dense motion trajectories, our method represents animation as a sequence of semantically meaningful keyframes in an interpretable ARKit-based facial control space. A language-driven model leverages large language model (LLM) priors to generate keyframes that align with contextual text descriptions and emotion cues. To support this formulation, we construct a multimodal dataset comprising 2,100 expression scripts paired with monocular videos, per-frame ARKit coefficients, and manually annotated semantic keyframes. Experiments show that incorporating semantic keyframe supervision and language priors significantly improves expression fidelity and semantic alignment compared to methods that do not use facial action semantics.

面部动画语言生成关键帧ARKit

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。