arXiv:2606.15848cs.CV2026-06

通过面部动作单元实现音频驱动3D人脸的精准表情控制。

EmoZone-Talker: Regional Semantic Control of Audio-Driven 3DGS Talking Heads via Facial Action Units

论文配图:EmoZone-Talker: Regional Semantic Control of Audio-Driven 3DGS Talking Heads via Facial Action Units
图 1 · 摘自论文原文
  • 用解耦区域约束和注意力偏置,分离语音与表情信号影响。
  • 引入独立通道时序编码器,提升表情动态连贯性。
  • 适合需要精细表情编辑的虚拟人、影视动画应用。

3D高斯点云(3DGS)在高保真人脸说话头生成中展现出巨大潜力。然而,由于语音驱动的面部动态与显式表情信号之间存在固有冲突,实现细粒度、可解释且可编辑的表情控制仍面临根本挑战。现有方法依赖隐式多模态融合,导致空间纠缠和时间不稳定。本文提出EmoZone-Talker,将音频驱动的面部动画重构为跨模态冲突下的结构化时空协调问题。我们引入带优先注意力偏置的协同区域(SZ-PAB),通过解剖先验指导的区域约束显式解耦模态贡献;并设计通道独立时序动作单元编码器(CIT-AE),建模时间上一致的动作单元动态。将这些表示集成到3D高斯变形中,实现了对表情的精确可解释控制。大量实验表明,该方法在上半脸准确性与时间连贯性方面均有显著提升,同时保持高渲染质量和精准口型同步。代码将公开以促进复现与进一步研究。

原文摘要 · Abstract (English)

3D Gaussian Splatting (3DGS) has shown strong potential for high-fidelity talking head synthesis. However, enabling fine-grained, interpretable, and editable facial expression control remains fundamentally challenging due to intrinsic conflicts between speech-driven facial dynamics and explicit expression signals. Existing methods rely on implicit multimodal fusion, leading to spatial entanglement and temporal instability. We present EmoZone-Talker, a novel framework that reformulates audio-driven facial animation as a structured spatial-temporal coordination problem under cross-modal conflicts. Our approach introduces an explicit spatial disentanglement and temporal dynamics modeling of facial motion. Specifically, we propose Synergy Zones with Prioritized Attention Bias (SZ-PAB) to explicitly decouple modality contributions via region-wise constraints guided by anatomical priors, and a Channel-Independent Temporal AU Encoder (CIT-AE) to model temporally coherent AU dynamics. By integrating these representations into 3D Gaussian deformation, EmoZone-Talker enables precise and interpretable control over facial expressions. Extensive experiments demonstrate that our method improves expression controllability and realism, with notable gains in upper-face accuracy and temporal coherence, while preserving high rendering quality and accurate lip synchronization. Code will be publicly released to facilitate reproducibility and further research.

3D高斯表情控制语音驱动动作单元

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。