实时生成对话角色的语音与口型动作,让虚拟形象自然对谈。
FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars

- 采用滚动流匹配与跨注意力机制,同步生成语音与面部动作。
- 在流式传输下实现高精度唇动同步与动作保真度。
- 适合虚拟主播、交互式数字人等需要真实对话体验的应用。
自然的人际对话需要实时语音生成与同步的面部动作。现有系统仅部分解决此问题:纯语音全双工模型可实时生成语音,但不产生面部动作;而音频驱动的面部动作模型则基于已有音频生成动作,无法在线联合生成语音与动作。为弥合这一差距,我们首次形式化了全双工联合语音-面部动作生成任务,即每一步同时生成语音标记与面部动作标记。在此基础上,提出FacePlex——一个统一的流式框架,包含两个核心组件:首先,滚动流匹配(Rolling Flow Matching)通过每步提交新动作帧,将流匹配适配于在线动作生成;其次,滚动交叉注意力(Rolling Cross-Attention)将流式音频队列与动作队列耦合,使语音与面部动作在生成过程中相互条件。通过大量实验、消融研究及用户测试,我们验证了FacePlex可在流式约束下实现全双工联合语音-面部动作生成,并在唇同步质量与动作保真度上优于音频驱动基线模型。
原文摘要 · Abstract (English)
Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. Existing systems only partially address this problem: speech-only full-duplex models can generate speech in real time but do not produce facial motion, while audio-driven facial motion models animate a face from already available audio rather than jointly generating speech and motion online. To bridge this gap, we first formalize full-duplex joint speech-facial motion generation, where speech tokens and facial motion tokens are produced together every step. Building on this formulation, we propose FacePlex, a unified streaming framework with two key components. First, Rolling Flow Matching adapts flow matching to online motion generation by committing new motion frames at each streaming step. Second, Rolling Cross-Attention couples the streaming audio queue with the motion queue, allowing speech and facial motion to condition each other as generation progresses. Through extensive experiments, ablation studies, and a user study, we show that FacePlex enables full-duplex joint speech-facial motion generation under online streaming constraints, while achieving stronger lip-sync quality and motion fidelity than audio-driven facial motion baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。