仅用文字生成逼真会动的虚拟人,说话与表情同步
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
- 双扩散变压器并行生成语音与视觉,中间有高速连接确保同步
- 从文本直接生成带口型、表情和头部动作的4D虚拟人
- 适合做对话式虚拟助手或数字人内容创作
我们提出AV-Flow,一种仅需文本输入即可生成逼真4D动态虚拟人的音视频生成模型。与以往依赖已有语音信号的工作不同,本方法联合生成语音与视觉内容。实验表明,该模型能生成自然的人类级语音、同步的口型动作、生动的表情与头部姿态,全部由文本字符驱动。核心架构采用两个并行的扩散变压器,通过中间高速连接实现音视频模态间的有效通信,从而保证语调与面部动态(如眉毛运动)的协调。模型采用流匹配训练策略,实现丰富表现力与快速推理。在双人对话场景中,AV-Flow可生成持续响应的虚拟人,实时感知并回应用户音视频输入。大量实验证明,该方法在生成逼真4D虚拟人方面优于现有工作。
原文摘要 · Abstract (English)
We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We demonstrate human-like speech synthesis, synchronized lip motion, lively facial expressions and head pose; all generated from just text characters. The core premise of our approach lies in the architecture of our two parallel diffusion transformers. Intermediate highway connections ensure communication between the audio and visual modalities, and thus, synchronized speech intonation and facial dynamics (e.g., eyebrow motion). Our model is trained with flow matching, leading to expressive results and fast inference. In case of dyadic conversations, AV-Flow produces an always-on avatar, that actively listens and reacts to the audio-visual input of a user. Through extensive experiments, we show that our method outperforms prior work, synthesizing natural-looking 4D talking avatars. Project page: https://aggelinacha.github.io/AV-Flow/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。