arXiv:2503.22605cs.GRcs.CV2025-03被引 3

用音频分解平面实现实时高精度人脸口型动画生成

Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head Synthesis

  • 将4维动态人脸建模分解为音视频独立的平面结构,提升表达效率
  • 在自驱动和跨驱动场景下均达到顶尖视觉质量与实时性能
  • 通过音频引导注意力机制聚焦嘴部区域,增强语音驱动动画精度

说话头合成是计算机图形学与多媒体领域的研究热点,但现有方法难以兼顾生成质量与计算效率,尤其在实时条件下。本文提出一种新框架,结合高斯点阵与结构化音频分解平面(Audio-Plane),实现高质量、音画同步且实时的说话头生成。传统方法需使用包含三维空间与时间轴的4D体积表示,但直接存储和处理密集4D网格会带来极高内存与计算开销,且难以扩展至长时长。为此,我们把4D体积分解为音视频独立与依赖的多个空间平面,构建紧凑可解释的音频平面(Audio-Plane)表示,显著提升模型对语音驱动复杂唇部动作的捕捉能力。为进一步增强局部运动建模,引入基于区域感知调制的音频引导显著性点阵机制,自适应强化如口部等高动态区域的关注度,使模型聚焦关键信息。大量实验表明,该方法在自驱动与跨驱动设置下均实现业界领先视觉质量、精准音唇同步与实时性能,优于现有2D与3D范式方法。

原文摘要 · Abstract (English)

Talking head synthesis has emerged as a prominent research topic in computer graphics and multimedia, yet most existing methods often struggle to strike a balance between generation quality and computational efficiency, particularly under real-time constraints. In this paper, we propose a novel framework that integrates Gaussian Splatting with a structured Audio Factorization Plane (Audio-Plane) to enable high-quality, audio-synchronized, and real-time talking head generation. For modeling a dynamic talking head, a 4D volume representation, which consists of three axes in 3D space and one temporal axis aligned with audio progression, is typically required. However, directly storing and processing a dense 4D grid is impractical due to the high memory and computation cost, and lack of scalability for longer durations. We address this challenge by decomposing the 4D volume representation into a set of audio-independent spatial planes and audio-dependent planes, forming a compact and interpretable representation for talking head modeling that we refer to as the Audio-Plane. This factorized design allows for efficient and fine-grained audio-aware spatial encoding, and significantly enhances the model's ability to capture complex lip dynamics driven by speech signals. To further improve region-specific motion modeling, we introduce an audio-guided saliency splatting mechanism based on region-aware modulation, which adaptively emphasizes highly dynamic regions such as the mouth area. This allows the model to focus its learning capacity on where it matters most for accurate speech-driven animation. Extensive experiments on both the self-driven and the cross-driven settings demonstrate that our method achieves state-of-the-art visual quality, precise audio-lip synchronization, and real-time performance, outperforming prior approaches across both 2D- and 3D-based paradigms.

说话头生成高斯点阵音频驱动实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。