提升情绪化3D说话人脸合成的精准度与一致性
Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation
- 分离音频内容与个性特征,实现情绪对齐与个性化
- 通过多模态注意力和4D高斯编码,精确捕捉微表情变化
- 根据视图不确定性自适应融合,提升渲染质量
情绪化说话人脸合成在多媒体与信号处理中至关重要,但现有3D方法存在两大挑战:音频-视觉情绪对齐差,表现为难以提取音频情绪且无法精确控制情绪微表情;以及通用多视角融合策略忽视不确定性与特征质量差异,影响生成质量。本文提出UA-3DTalk,包含三个核心模块:先验提取模块将音频解耦为内容同步特征与个性化互补特征;情绪蒸馏模块引入多模态注意力加权融合与多分辨率码本的4D高斯编码,实现细粒度音频情绪提取与微表情精确控制;不确定性变形模块通过不确定性块估计视图特定的随机性(输入噪声)与认知性(模型参数)不确定性,实现自适应多视角融合,并采用多头解码器优化高斯基元,克服均匀加权融合的局限。在常规与情绪数据集上的大量实验表明,相较于DEGSTalk与EDTalk等前沿方法,UA-3DTalk在情绪对齐上提升5.2%(E-FID)、唇音同步提升3.1%(SyncC),渲染质量提升0.015(LPIPS)。
原文摘要 · Abstract (English)
Emotional Talking Face synthesis is pivotal in multimedia and signal processing, yet existing 3D methods suffer from two critical challenges: poor audio-vision emotion alignment, manifested as difficult audio emotion extraction and inadequate control over emotional micro-expressions; and a one-size-fits-all multi-view fusion strategy that overlooks uncertainty and feature quality differences, undermining rendering quality. We propose UA-3DTalk, Uncertainty-Aware 3D Emotional Talking Face Synthesis with emotion prior distillation, which has three core modules: the Prior Extraction module disentangles audio into content-synchronized features for alignment and person-specific complementary features for individualization; the Emotion Distillation module introduces a multi-modal attention-weighted fusion mechanism and 4D Gaussian encoding with multi-resolution code-books, enabling fine-grained audio emotion extraction and precise control of emotional micro-expressions; the Uncertainty-based Deformation deploys uncertainty blocks to estimate view-specific aleatoric (input noise) and epistemic (model parameters) uncertainty, realizing adaptive multi-view fusion and incorporating a multi-head decoder for Gaussian primitive optimization to mitigate the limitations of uniform-weight fusion. Extensive experiments on regular and emotional datasets show UA-3DTalk outperforms state-of-the-art methods like DEGSTalk and EDTalk by 5.2% in E-FID for emotion alignment, 3.1% in SyncC for lip synchronization, and 0.015 in LPIPS for rendering quality. Project page: https://mrask999.github.io/UA-3DTalk
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。