首个可通用的音频驱动写实人脸生成方法,能精准还原口型与表情细节。
Audio-Driven Universal Gaussian Head Avatars
- 用通用头部先验模型结合音频直接生成表情,同时建模几何与外观变化。
- 在LipSync、图像质量和感知真实度上均超越现有几何类方法。
- 适合需要高保真音视频同步人脸生成的场景,如虚拟主播或数字人。
我们提出首个音频驱动的通用写实人脸生成方法,结合无身份依赖的语音模型与新型通用头部先验(UHAP)。UHAP 在跨身份多视角视频上训练,并以中性扫描数据为监督,能高保真捕捉个体特异性细节。与以往仅将音频映射到几何变形的方法不同,我们的通用语音模型直接将原始音频输入映射至UHAP潜在表达空间,该空间同时编码几何与外观变化。为实现对新主体的高效个性化,采用单目编码器轻量级回归帧间动态表达变化,使后续微调阶段专注捕捉主体全局外观与几何。通过UHAP解码音频驱动的表达码,生成具有精确口型同步及细微表情细节(如眉毛动作、视线转移、口腔内部外观与运动)的高真实感头像。大量评估表明,本方法是首个可泛化且能建模精细外观的音频驱动头像模型,在唇同步准确率、图像质量与感知真实度指标上均优于现有几何类方法。
原文摘要 · Abstract (English)
We introduce the first method for audio-driven universal photorealistic avatar synthesis, combining a person-agnostic speech model with our novel Universal Head Avatar Prior (UHAP). UHAP is trained on cross-identity multi-view videos. In particular, our UHAP is supervised with neutral scan data, enabling it to capture the identity-specific details at high fidelity. In contrast to previous approaches, which predominantly map audio features to geometric deformations only while ignoring audio-dependent appearance variations, our universal speech model directly maps raw audio inputs into the UHAP latent expression space. This expression space inherently encodes, both, geometric and appearance variations. For efficient personalization to new subjects, we employ a monocular encoder, which enables lightweight regression of dynamic expression variations across video frames. By accounting for these expression-dependent changes, it enables the subsequent model fine-tuning stage to focus exclusively on capturing the subject's global appearance and geometry. Decoding these audio-driven expression codes via UHAP generates highly realistic avatars with precise lip synchronization and nuanced expressive details, such as eyebrow movement, gaze shifts, and realistic mouth interior appearance as well as motion. Extensive evaluations demonstrate that our method is not only the first generalizable audio-driven avatar model that can account for detailed appearance modeling and rendering, but it also outperforms competing (geometry-only) methods across metrics measuring lip-sync accuracy, quantitative image quality, and perceptual realism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。