用音频生成情感丰富、长时序逼真的虚拟人脸表演
X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio
- 分两阶段生成:先预测情绪化面部动作,再合成高清视频
- 支持无限长度情感表达,无误差累积,保持语音与表情同步
- 适合影视特效、虚拟主播等需要自然情绪表达的场景
我们提出X-Actor,一种全新的音频驱动人脸动画框架,仅需一张参考图像和一段音频输入,即可生成逼真且富有情感表达的说话头视频。与以往侧重唇形同步和短时视觉保真度的方法不同,X-Actor实现了高质量、长时序的人脸表演,能捕捉随语音节奏与内容动态变化的细腻情绪。核心是两阶段解耦生成流程:第一阶段为音频条件的自回归扩散模型,在长时间上下文中预测去身份化的面部运动潜在表示;第二阶段为基于扩散模型的视频生成模块,将这些运动转换为高保真视频动画。通过在解耦的面部运动潜在空间中操作,模型借助扩散强化训练范式,有效捕捉音频与面部动态间的长程相关性,实现无误差积累的无限长度情感化动作预测。大量实验表明,X-Actor生成的表演具有电影级表现力,显著超越传统说话头动画,在长时序、音频驱动的情感人脸表演任务上达到当前最优效果。
原文摘要 · Abstract (English)
We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that emphasize lip synchronization and short-range visual fidelity in constrained speaking scenarios, X-Actor enables actor-quality, long-form portrait performance capturing nuanced, dynamically evolving emotions that flow coherently with the rhythm and content of speech. Central to our approach is a two-stage decoupled generation pipeline: an audio-conditioned autoregressive diffusion model that predicts expressive yet identity-agnostic facial motion latent tokens within a long temporal context window, followed by a diffusion-based video synthesis module that translates these motions into high-fidelity video animations. By operating in a compact facial motion latent space decoupled from visual and identity cues, our autoregressive diffusion model effectively captures long-range correlations between audio and facial dynamics through a diffusion-forcing training paradigm, enabling infinite-length emotionally-rich motion prediction without error accumulation. Extensive experiments demonstrate that X-Actor produces compelling, cinematic-style performances that go beyond standard talking head animations and achieves state-of-the-art results in long-range, audio-driven emotional portrait acting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。