用3分钟视频生成实时逼真人脸动画,速度与质量双优。
EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
- 基于3D高斯点云,分静态初始化与音频驱动变形两阶段构建
- 仅需3-5分钟训练,渲染质量与唇同步精度达顶尖水平
- 适合需要低延迟、高保真的直播/虚拟人应用
本文提出EGSTalker,一种基于3D高斯点云(3DGS)的实时音频驱动人脸生成框架。该框架仅需3-5分钟训练视频即可生成高质量面部动画。系统包含两个关键阶段:静态高斯初始化与音频驱动变形。第一阶段利用多分辨率哈希三平面和柯尔莫哥洛夫-阿诺德网络(KAN)提取空间特征并构建紧凑的3D高斯表示。第二阶段提出高效空间-音频注意力(ESAA)模块融合音频与空间线索,同时由KAN预测对应的高斯形变。大量实验表明,EGSTalker在渲染质量和唇同步准确率上达到当前最优水平,且推理速度显著优于现有方法。结果凸显其在实时多媒体应用中的潜力。
原文摘要 · Abstract (English)
This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3-5 minutes of training video to synthesize high-quality facial animations. The framework comprises two key stages: static Gaussian initialization and audio-driven deformation. In the first stage, a multi-resolution hash triplane and a Kolmogorov-Arnold Network (KAN) are used to extract spatial features and construct a compact 3D Gaussian representation. In the second stage, we propose an Efficient Spatial-Audio Attention (ESAA) module to fuse audio and spatial cues, while KAN predicts the corresponding Gaussian deformations. Extensive experiments demonstrate that EGSTalker achieves rendering quality and lip-sync accuracy comparable to state-of-the-art methods, while significantly outperforming them in inference speed. These results highlight EGSTalker's potential for real-time multimedia applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。