arXiv:2411.16331cs.MMcs.CV2024-11CVPR被引 81

用全局音频感知提升人脸动画自然度,告别依赖视觉的僵硬表现。

Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

  • 分离音频的片段内与片段间感知,用语音语调和节奏引导表情动作
  • 在多个数据集上实现更精准的口型同步与更连贯的动作序列
  • 适合追求真实感语音驱动动画的研究者与开发者

现有说话人脸生成方法多依赖视觉与空间信息稳定动作,导致自然度下降。本文提出Sonic新范式,聚焦全局音频感知:通过上下文增强的音频学习提取语音语调与节奏作为表情和嘴部动作先验;采用运动解耦控制器独立控制头部与表情运动;并引入时间感知的位置偏移融合机制,利用连续的时间窗口融合跨片段音频信息以实现长音频推理。大量实验表明,该方法在视频质量、时序一致性、口型同步精度和动作多样性上均优于现有最先进方法。

原文摘要 · Abstract (English)

The study of talking face generation mainly explores the intricacies of synchronizing facial movements and crafting visually appealing, temporally-coherent animations. However, due to the limited exploration of global audio perception, current approaches predominantly employ auxiliary visual and spatial knowledge to stabilize the movements, which often results in the deterioration of the naturalness and temporal inconsistencies.Considering the essence of audio-driven animation, the audio signal serves as the ideal and unique priors to adjust facial expressions and lip movements, without resorting to interference of any visual signals. Based on this motivation, we propose a novel paradigm, dubbed as Sonic, to {s}hift f{o}cus on the exploration of global audio per{c}ept{i}o{n}.To effectively leverage global audio knowledge, we disentangle it into intra- and inter-clip audio perception and collaborate with both aspects to enhance overall perception.For the intra-clip audio perception, 1). \textbf{Context-enhanced audio learning}, in which long-range intra-clip temporal audio knowledge is extracted to provide facial expression and lip motion priors implicitly expressed as the tone and speed of speech. 2). \textbf{Motion-decoupled controller}, in which the motion of the head and expression movement are disentangled and independently controlled by intra-audio clips. Most importantly, for inter-clip audio perception, as a bridge to connect the intra-clips to achieve the global perception, \textbf{Time-aware position shift fusion}, in which the global inter-clip audio information is considered and fused for long-audio inference via through consecutively time-aware shifted windows. Extensive experiments demonstrate that the novel audio-driven paradigm outperform existing SOTA methodologies in terms of video quality, temporally consistency, lip synchronization precision, and motion diversity.

语音驱动人脸动画音频感知动作解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。