arXiv:2503.01715cs.CVcs.AI2025-03CVPR被引 12

用关键帧插值实现长时自然语音驱动面部动画

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

  • 分两阶段生成:先低频生成关键帧,再插值填补中间帧
  • 在长达10秒以上序列中保持身份一致且动作连贯
  • 支持笑声、叹息等非语音发声,适合影视动画场景

当前语音驱动面部动画方法在短视频上表现优异,但在长时间序列中易出现误差累积和身份漂移。现有方法通过外部空间控制提升长期一致性,但牺牲了动作自然性。本文提出KeyFace,一种基于扩散模型的两阶段框架:第一阶段以低帧率生成关键帧,结合音频输入与身份帧捕捉长时间内的核心表情与动作;第二阶段通过插值模型填充关键帧间的空隙,确保过渡平滑与时间连贯性。为增强真实感,引入连续情绪表征,并处理各类非语音发声(NSVs),如笑声与叹息。此外,提出两个新评估指标,用于衡量唇同步与NSV生成效果。实验表明,KeyFace在超过10秒的长序列上优于现有最优方法,成功生成自然、连贯的面部动画,准确涵盖非语音发声与连续情绪变化。

原文摘要 · Abstract (English)

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. Existing methods attempt to mitigate this through external spatial control, increasing long-term consistency but compromising the naturalness of motion. We propose KeyFace, a novel two-stage diffusion-based framework, to address these issues. In the first stage, keyframes are generated at a low frame rate, conditioned on audio input and an identity frame, to capture essential facial expressions and movements over extended periods of time. In the second stage, an interpolation model fills in the gaps between keyframes, ensuring smooth transitions and temporal coherence. To further enhance realism, we incorporate continuous emotion representations and handle a wide range of non-speech vocalizations (NSVs), such as laughter and sighs. We also introduce two new evaluation metrics for assessing lip synchronization and NSV generation. Experimental results show that KeyFace outperforms state-of-the-art methods in generating natural, coherent facial animations over extended durations, successfully encompassing NSVs and continuous emotions.

语音驱动面部动画扩散模型长序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。