arXiv:2504.05746cs.CV2025-04中稿 · TMM 2025被引 3

利用音视频时序相关性提升单张图像说话头动画的细节表现

Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation

  • 设计时序音视频相关嵌入框架,对齐音频与视频帧间的时序关系
  • 在多个基准数据集上显著优于现有方法,生成更自然的微表情变化
  • 适合需要高保真说话头动画的虚拟主播、数字人场景

音视频驱动的一次性说话头动画(ADOS-THA)的核心挑战在于捕捉相邻视频帧间细微且难以察觉的变化。本质上,相邻音频片段的时序关系与对应相邻视频帧的时序关系高度相关,可为说话头动画提供关键补充信息。本文提出一种新型时序音视频相关嵌入(TAVCE)框架,通过学习并整合音视频时序相关性来增强特征表示并正则化生成结果。具体而言,该框架首先学习一个音视频时序相关度量,确保相邻音频片段的时序关系与对应视频帧的时序关系对齐;由于时序音频关系包含关于视觉帧的对齐信息,我们通过简单有效的通道注意力机制将其融入特征学习以生成更具代表性特征;训练过程中,还利用对齐的相关性作为额外监督目标指导视频帧生成。我们在多个公开基准数据集(包括HDTF、LRW、VoxCeleb1和VoxCeleb2)上进行了大量实验,验证了其相对于现有领先算法的优越性。

原文摘要 · Abstract (English)

The paramount challenge in audio-driven One-shot Talking Head Animation (ADOS-THA) lies in capturing subtle imperceptible changes between adjacent video frames. Inherently, the temporal relationship of adjacent audio clips is highly correlated with that of the corresponding adjacent video frames, offering supplementary information that can be pivotal for guiding and supervising talking head animations. In this work, we propose to learn audio-visual correlations and integrate the correlations to help enhance feature representation and regularize final generation by a novel Temporal Audio-Visual Correlation Embedding (TAVCE) framework. Specifically, it first learns an audio-visual temporal correlation metric, ensuring the temporal audio relationships of adjacent clips are aligned with the temporal visual relationships of corresponding adjacent video frames. Since the temporal audio relationship contains aligned information about the visual frame, we first integrate it to guide learning more representative features via a simple yet effective channel attention mechanism. During training, we also use the alignment correlations as an additional objective to supervise generating visual frames. We conduct extensive experiments on several publicly available benchmarks (i.e., HDTF, LRW, VoxCeleb1, and VoxCeleb2) to demonstrate its superiority over existing leading algorithms.

说话头动画音视频对齐生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。