用文本控制多视角生动人脸动画,实现动作表情精准可控。
MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait Animation
- 以FLAME参数为中间表示,分离建模运动与情绪
- 支持多视角生成,文本驱动下动作表情更自然
- 兼容文本/语音/视频多种输入,适合交互式创作
近期人脸动画方法在唇形同步方面取得显著进展,但普遍缺乏对头部动作和面部表情的显式控制,且无法生成多视角视频,导致动画可控性与表现力不足。此外,尽管文本驱动方式更友好,相关研究仍较薄弱。本文提出首个两阶段文本驱动框架MVPortrait(多视角生动人脸动画),通过将FLAME作为中间表示,将面部动作、表情及视角变换嵌入其参数空间。第一阶段基于文本输入分别训练FLAME运动与情绪扩散模型;第二阶段训练多视角视频生成模型,以参考人脸图像和第一阶段生成的多视角FLAME渲染序列为条件。实验表明,MVPortrait在动作与情绪控制、视角一致性上均优于现有方法。更重要的是,借助FLAME桥梁,该框架首次实现对文本、语音、视频等信号的统一可控生成。
原文摘要 · Abstract (English)
Recent portrait animation methods have made significant strides in generating realistic lip synchronization. However, they often lack explicit control over head movements and facial expressions, and cannot produce videos from multiple viewpoints, resulting in less controllable and expressive animations. Moreover, text-guided portrait animation remains underexplored, despite its user-friendly nature. We present a novel two-stage text-guided framework, MVPortrait (Multi-view Vivid Portrait), to generate expressive multi-view portrait animations that faithfully capture the described motion and emotion. MVPortrait is the first to introduce FLAME as an intermediate representation, effectively embedding facial movements, expressions, and view transformations within its parameter space. In the first stage, we separately train the FLAME motion and emotion diffusion models based on text input. In the second stage, we train a multi-view video generation model conditioned on a reference portrait image and multi-view FLAME rendering sequences from the first stage. Experimental results exhibit that MVPortrait outperforms existing methods in terms of motion and emotion control, as well as view consistency. Furthermore, by leveraging FLAME as a bridge, MVPortrait becomes the first controllable portrait animation framework that is compatible with text, speech, and video as driving signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。